Alina Schanz

Worked with · 09 of 12AI trainingprimeintellect.ai

Prime Intellect

I collaborated with the Prime Intellect team on research around agentic reinforcement learning, training environments and the infrastructure needed to evaluate agents that operate over longer horizons.

Relationship
Worked with
Company
Infrastructure for training AI models and the environments agents learn in
Format
Research collaboration
Focus
Agentic reinforcement learning and training environments
Covered
Environments, harnesses and runtimes, verification, sandboxes, long-horizon tasks, browser agents

The environment

A large part of my work focused on the environment itself. With agentic RL, the model is only one part of the system. The environment defines what the agent can see, which tools it can use, what state it can change and how its actions are scored. Small choices in that setup can change the behavior the model learns.

Fig. 1What the environment decides

  1. 01What the agent can see
  2. 02Which tools it can use
  3. 03What state it can change
  4. 04How its actions are scored

Three layers

I spent time looking at how tasks, agent harnesses and execution runtimes should be separated. A coding task might stay the same while the agent solving it changes from a simple ReAct loop to a CLI coding agent with subagents, compaction and its own tool logic. Keeping those layers separate makes it easier to tell whether an improvement came from the model, the harness or the environment around it.

Fig. 2Kept apart on purpose

  1. TaskThe coding task stays the same
  2. HarnessA simple ReAct loopOr a CLI coding agent with subagents, compaction and its own tool logic
  3. RuntimeWhere the work is executed
With the layers separate, an improvement can be traced to the model, the harness or the environment around it.

Verification

Verification was another major part of the work. Agent training depends on reward, so the quality of the verifier matters as much as the task itself. I looked at cases where a model can satisfy a narrow grader without solving the underlying problem, as well as the opposite case where a correct solution is rejected because the test encodes one expected implementation too rigidly.

Fig. 3Two ways a verifier fails

A

False positive

  • A narrow grader is satisfied
  • The underlying problem is not solved
B

False negative

  • The solution is correct
  • The test expects one implementation too rigidly

At RL scale

That becomes more serious at RL scale. A weak verifier does not only produce a bad evaluation score. It becomes part of the training signal, which means the model can actively learn the wrong behavior. I was interested in reward hacking, false positives and false negatives, and in separating the environment the agent can manipulate from the system responsible for grading its work.

Fig. 4Where a weak verifier ends up

  1. 01TaskWhat the agent is asked
  2. 02AgentActs in an environment it can change
  3. 03VerifierGrades from outside its reach
  4. 04RewardRight or wrong
  5. 05TrainingWhat the model learns next

Sandboxes

Sandboxes were part of the same problem. Agents that execute code, operate browsers or modify files need real stateful environments, but training can require thousands of those environments to run concurrently. I looked at isolation, reproducibility, startup latency and how much environment state should survive between steps of a rollout.

Fig. 5Thousands of environments at once

  1. Isolation
  2. Reproducibility
  3. Startup latency
  4. State between steps

Long horizons

Long-horizon tasks made those tradeoffs more visible. Once an agent works for hours instead of answering a single prompt, recovery, state management, tool failures and partial progress become part of the evaluation. A benchmark can say that an agent failed without telling you whether the failure came from reasoning, execution infrastructure or a brittle verifier.

Fig. 6One failed run, three possible causes

A

Reasoning

B

Execution infrastructure

C

A brittle verifier

Browsers and computers

I also looked at browser and computer-use agents, where the environment is much less predictable than a static benchmark. Live pages change, sessions expire, layouts move and external systems fail. Training against that kind of environment requires a different standard for both reproducibility and task verification.

Measurable

The broader question throughout the collaboration was how to make agent improvement measurable. Better benchmark scores are useful, but I was more interested in whether an agent could complete real tasks reliably, recover when something went wrong and improve without learning shortcuts in the reward system.

What I took from it

That work changed how I think about agentic systems. The model matters, but a large part of progress comes from the infrastructure around it: the tasks it receives, the tools it can use, the environment it operates in and the mechanism that decides whether its work was actually correct.

Fig. 7Around the model

  1. The tasks it receives
  2. The tools it can use
  3. The environment it operates in
  4. What decides whether its work was correct

Materials