Worked with · 09 of 12AI trainingprimeintellect.ai
Prime Intellect
I collaborated with the Prime Intellect team on research around agentic reinforcement learning, training environments and the infrastructure needed to evaluate agents that operate over longer horizons.
The environment
A large part of my work focused on the environment itself. With agentic RL, the model is only one part of the system. The environment defines what the agent can see, which tools it can use, what state it can change and how its actions are scored. Small choices in that setup can change the behavior the model learns.
Fig. 1What the environment decides
- 01What the agent can see
- 02Which tools it can use
- 03What state it can change
- 04How its actions are scored
Three layers
I spent time looking at how tasks, agent harnesses and execution runtimes should be separated. A coding task might stay the same while the agent solving it changes from a simple ReAct loop to a CLI coding agent with subagents, compaction and its own tool logic. Keeping those layers separate makes it easier to tell whether an improvement came from the model, the harness or the environment around it.
Fig. 2Kept apart on purpose
- TaskThe coding task stays the same
- HarnessA simple ReAct loopOr a CLI coding agent with subagents, compaction and its own tool logic
- RuntimeWhere the work is executed
Verification
Verification was another major part of the work. Agent training depends on reward, so the quality of the verifier matters as much as the task itself. I looked at cases where a model can satisfy a narrow grader without solving the underlying problem, as well as the opposite case where a correct solution is rejected because the test encodes one expected implementation too rigidly.
Fig. 3Two ways a verifier fails
False positive
- A narrow grader is satisfied
- The underlying problem is not solved
False negative
- The solution is correct
- The test expects one implementation too rigidly
At RL scale
That becomes more serious at RL scale. A weak verifier does not only produce a bad evaluation score. It becomes part of the training signal, which means the model can actively learn the wrong behavior. I was interested in reward hacking, false positives and false negatives, and in separating the environment the agent can manipulate from the system responsible for grading its work.
Fig. 4Where a weak verifier ends up
- 01TaskWhat the agent is asked
- 02AgentActs in an environment it can change
- 03VerifierGrades from outside its reach
- 04RewardRight or wrong
- 05TrainingWhat the model learns next
Sandboxes
Sandboxes were part of the same problem. Agents that execute code, operate browsers or modify files need real stateful environments, but training can require thousands of those environments to run concurrently. I looked at isolation, reproducibility, startup latency and how much environment state should survive between steps of a rollout.
Fig. 5Thousands of environments at once
- Isolation
- Reproducibility
- Startup latency
- State between steps
Long horizons
Long-horizon tasks made those tradeoffs more visible. Once an agent works for hours instead of answering a single prompt, recovery, state management, tool failures and partial progress become part of the evaluation. A benchmark can say that an agent failed without telling you whether the failure came from reasoning, execution infrastructure or a brittle verifier.
Fig. 6One failed run, three possible causes
Reasoning
Execution infrastructure
A brittle verifier
Browsers and computers
I also looked at browser and computer-use agents, where the environment is much less predictable than a static benchmark. Live pages change, sessions expire, layouts move and external systems fail. Training against that kind of environment requires a different standard for both reproducibility and task verification.
Measurable
The broader question throughout the collaboration was how to make agent improvement measurable. Better benchmark scores are useful, but I was more interested in whether an agent could complete real tasks reliably, recover when something went wrong and improve without learning shortcuts in the reward system.
What I took from it
That work changed how I think about agentic systems. The model matters, but a large part of progress comes from the infrastructure around it: the tasks it receives, the tools it can use, the environment it operates in and the mechanism that decides whether its work was actually correct.
Fig. 7Around the model
- The tasks it receives
- The tools it can use
- The environment it operates in
- What decides whether its work was correct
Materials
Public pages about the company and the programs around this work.
Screenshot, Oct 5, 2026
Prime Intellect blog · July 2026
verifiers v1: decomposing tasksets and harnesses
Decomposing tasksets and harnesses for agentic RL and evaluations.
Further reading