Forward-deployed AI engineering
Agents that finish the work.
We train models against your real workflows — long, multi-step, mostly in interfaces that were never meant for machines — and stay embedded until they run in production.
What we bring
Careers spent on large-scale training: SFT and RL runs at frontier scale, the environments they train against, and the harnesses that carry a model into production.
4 phases
Evaluate, collect, train, deploy — one owner
100s
Steps per task — the horizon we train for
1 team
Environments, training and deployment together
The demo was never the problem.
Three things stop a capable model from becoming a reliable system. All three are training problems wearing an engineering costume.
Prompting plateaus
A general model reaches seventy percent of your workflow and stalls. The remaining thirty percent is where the value sits, and no prompt will buy it.
Nobody owns the loop
Environments, data, training and deployment sit in four different teams. The loop never closes, so the model never improves after launch.
Long horizons compound down
Ninety-five percent per step is well under one percent across a hundred steps. Reliability at length has to be trained in — recovery, verification, knowing when to stop.
One loop, four phases.
Every engagement runs the same shape. Fixed scope at each step, and a deliverable you can check.
01
Evaluate
We instrument the workflow and build the eval that actually predicts production behaviour.
Out: harness and honest baseline
02
Build the environment
Your workflow rebuilt as something a model can practise in, plus the demonstrations and traces to seed it.
Out: a trainable copy of the work
03
Train
SFT to teach the shape of the work, RL against the environment to make it survive drift, distillation to make it affordable.
Out: a model that beats baseline
04
Deploy
Rollout behind your guardrails, monitoring on the metrics you chose, and a retrain path we hand over.
Out: it runs, and it keeps improving
What we are unusually good at.
Six things, and they are the same six on every engagement.
Enterprise workflow agents
Agents that run a real business process end to end — legacy portals, internal tools, vendor consoles, anything with no API and no appetite for one.
Agent harnesses
Tool interfaces, memory, budgets and retries — plus traces good enough to debug a failure at step ninety.
RL environments
We rebuild your workflow as something a model can practise in — deterministic, resettable, and scored the way you would score a person.
RL at scale
Distributed rollouts, reward design and the training infrastructure that keeps a long run stable for weeks rather than hours.
Large-scale training
Pre-training and post-training at scale: SFT, reward modelling and distillation — on your weights, on open weights, or through a frontier model's tuning surface.
Evaluation
The unglamorous half. Harnesses that tell you the truth about a model before your customers do it for you.
Fixed scope. Named team.
No open-ended discovery, no bench staffing. You know the price and the deliverable before each phase starts.
A scoping week tells you whether the workflow is trainable. If it isn't, we say so and you keep the workflow map.
| Phase | Duration | What you get |
|---|---|---|
| Scoping | 1 week | Workflow map, feasibility read, fixed price for the pilot |
| Pilot | 6 weeks | Eval harness, environment, first trained candidate |
| Production | 12 weeks | Deployed model, monitoring, retrain loop, runbook |
| Standing team | Ongoing | Embedded engineers against quarterly capability targets |
The team
Frontier-scale training, pointed at one workflow.
We have spent our careers on large-scale model training — pre-training and post-training runs at frontier scale, reinforcement learning across thousands of parallel environments, and the agent systems built on top of them. We left to do that work one customer at a time.
Small on purpose. The people who train the model are the people on your call.

If you are an enterprise
You have a workflow that eats thousands of hours a year and every vendor has told you it is out of scope. That is the one we want.
If you are an AI lab
You have capability and no capacity. We take a customer end to end as your forward-deployed team, and hand the recipe back to you.