MachinoAI explainer / AI Agents
OpenForgeRL: Train Harness-native Agents in Any Environment
OpenForgeRL makes complex production-style agent harnesses trainable end-to-end by separating remote environment rollouts from standard RL training infrastructure.
01
Abstract
A useful way to interpret the contribution is that OpenForgeRL changes the unit of training from “model plus simplified tool loop” to “model operating through its real harness.” This matters because the harness determines what information is surfaced to the model and which actions are actually possible. The paper therefore evaluates not only final task success but also how training changes intermediate agent behavior.
02
Introduction
Agentic AI has moved beyond single-turn generation. Coding agents, browser agents, computer-use systems, and tool-using assistants now operate through software harnesses that decide how the model sees context, which tools are available, how subagents are invoked, and how actions are executed. The paper uses Claude Code, Codex, and OpenClaw as examples of this broader trend. These harnesses can substantially change the effective behavior of the same underlying model.\n\nThis creates a training problem. Conventional RL infrastructure generally assumes that the trainer controls model generation directly. A sophisticated harness instead owns the interaction loop: it may make multiple model calls, invoke tools, launch subagents, maintain hidden state, and interact with a dedicated environment. Reproducing all of that inside a training framework is expensive and often produces a simplified version of the deployment system.\n\nOpenForgeRL treats the harness as part of the environment rather than as something to be rewritten for training. Its design separates two concerns: the existing harness remains responsible for inference behavior, while the training system supplies model generations, records trajectories, calculates rewards, and updates the policy. Remote containers then provide the CPU, memory, browser, filesystem, or other resources required by each environment.\n\nThe result is a framework intended to reduce the gap between what an agent learns during training and what it encounters during deployment. The paper evaluates this idea across both text-based tool-use agents and multimodal GUI agents, making the work particularly relevant to production agent infrastructure.
03
Problem
The core problem is the train–deploy mismatch created by complex inference harnesses. A bare LLM can be trained with standard RL because the trainer can directly construct prompts, generate tokens, and collect rewards. A production agent is different: the harness may manage tools, subagents, context windows, skills, environment state, and multiple model calls. The trainer does not necessarily see or control the complete internal execution flow.\n\nA second problem is infrastructure isolation. Agent rollouts may require browsers, desktops, terminals, network access, filesystems, dedicated CPU and memory, or other environment-specific resources. Standard RL implementations commonly assume that rollouts execute close to the training process. At scale, this is inefficient or impossible for heavyweight environments.\n\nThe paper therefore asks whether an open training system can connect arbitrary agent harnesses and environments to existing RL infrastructure without reimplementing each harness. The desired interface should support stateful multi-turn execution, remote environments, asynchronous rollouts, reward collection, and trajectory reconstruction while remaining largely agnostic to the underlying RL algorithm.\n\nA production implication follows directly: if the harness is responsible for a large part of an agent’s behavior, improving only the model weights while ignoring the harness can leave a significant portion of the deployed system unoptimized.
The problem becomes especially visible in long-horizon tasks. A rollout may contain many model calls whose dependencies are hidden inside the harness. A training system that sees only a final answer can struggle to assign useful learning signals to the individual decisions that produced it. OpenForgeRL does not solve fine-grained credit assignment completely, but it makes the actual harness-level interactions available as a standard trajectory, which is an important prerequisite for studying the problem.
04
Background
The paper builds on two strands of prior work. The first is the emergence of agent inference harnesses. Software-engineering agents and general-purpose assistants increasingly use orchestration layers to manage tools, context, skills, and control flow. The second is open RL infrastructure such as veRL, Slime, OpenRLHF, and related systems. These frameworks provide scalable optimization algorithms, but their rollout assumptions are often simpler than those of full production harnesses.\n\nOpenForgeRL positions itself at the boundary between these layers. Rather than proposing a new language-model architecture or a new RL objective, it provides infrastructure for connecting the existing harness layer to the existing training layer. The paper also discusses concurrent work such as Polar, which similarly targets harness-level agentic RL but focuses on software-engineering tasks, whereas OpenForgeRL evaluates a broader mix of tool-use and GUI settings.\n\nThe important conceptual shift is to treat the inference harness as part of the agent being trained. The model policy still produces actions, but the effective policy is mediated by tools, control flow, and environment interfaces. Training against the same interfaces used at deployment can therefore expose behaviors that a simplified training loop may miss.
This framing also connects agent evaluation to infrastructure design. Two agents with identical base weights can behave differently when one has better tool interfaces, context management, or control flows. OpenForgeRL makes those differences experimentally visible because the same trained model can be evaluated through multiple harnesses. The paper’s harness-transfer experiments are therefore part of the scientific contribution rather than merely an implementation detail.
05
Methodology
OpenForgeRL has two central components. First, a lightweight proxy wraps the model inference server and intercepts model-generation requests coming from the real harness. Second, a Kubernetes orchestrator creates and manages remote rollout containers. The harness and environment run inside those containers, while the RL trainer remains separate.\n\nDuring a rollout, the proxy records the harness-level prompt-response interactions and the environment returns a terminal reward, typically task success. OpenForgeRL reconstructs these interactions into training trajectories that can be consumed by a standard RL framework such as veRL. This makes the method largely independent of the harness implementation and allows the environment definition to contain most of the harness-specific integration.
The remote-rollout design also addresses practical failure modes. Because individual environments can stall, OpenForgeRL uses wall-clock timeouts rather than relying only on a model turn limit. Failed rollouts caused by infrastructure or harness errors are discarded rather than treated as ordinary policy failures, avoiding misleading negative training signals from partial trajectories. The authors identify better credit assignment for these partial failures as future work.
Figure notes
Visual evidence
Formula notes
Mathematical details
Trajectory reward assignment
The terminal reward is propagated backward through the trajectory using a discount factor. The paper typically uses gamma = 1.0, so every step receives the same terminal reward signal.
a_t- -model action/response at step t
r_T- -terminal reward
r_t- -reward assigned to step t
tau- -reconstructed training trajectory
gamma- -discount factor
s_H_t- -harness-level state/input at step t
06
Architecture
The system architecture separates training from execution. An RL trainer and inference server sit outside the rollout environments. A proxy receives generation requests from remote harness instances and routes them to the inference server while recording the exchanged inputs and outputs. A Kubernetes-based orchestrator creates rollout pods, assigns resources, monitors them, and removes them when the rollout ends.\n\nThe environment can contain a CLI, browser, desktop, tool server, or other task-specific interface. Supporting a new harness or environment primarily requires changing the sandbox rather than rewriting the training loop. This separation is the key architectural contribution: the deployment harness remains intact while the optimization system treats its observable model calls and terminal rewards as a training interface.
The separation has a useful scaling property: training GPUs do not need to host every environment. Remote rollout pods can scale independently, while the training backend continues to process reconstructed trajectories. In a production implementation, this creates a natural boundary between GPU-intensive model optimization and CPU/network/display-intensive environment execution. It also introduces distributed-systems concerns such as scheduling, timeouts, resource isolation, networking, and observability.
07
Dataset
The task-generation process is important because agent RL is often limited by executable environments rather than by raw text examples. Each generated task is expected to have an instruction, environment, artifacts, and a verifier. The pipeline tests and repairs the task before it enters training, reducing the chance that a broken environment becomes a source of noisy reward. The paper reports only hundreds to a few thousand tasks, emphasizing the value of environment quality and rollout diversity rather than simply dataset scale.
08
Training
For Claw experiments, the backbone is Qwen3-30B-A3B-Thinking. SFT distills successful trajectories from MiniMax-M2.5 using three sampled rollouts per task. RL then continues from the SFT checkpoint using GRPO, with veRL as the training backend, Microsoft Azure for rollout containers, batch size 8, group size 8, and 8 B200 GPUs.\n\nFor GUI experiments, the backbone is Qwen3-VL-8B-Thinking and successful trajectories are distilled from Kimi-K2.5. The same GRPO, veRL, Azure, batch-size, group-size, and 8×B200 setup is used; screenshots are used as visual input during training and evaluation. The paper provides additional hyperparameters and training curves in the appendix.
The training setup therefore combines supervised imitation with reinforcement learning. SFT provides a starting policy from successful teacher trajectories; GRPO then optimizes behavior using rewards obtained from executing the agent in the environment. The important point is that the RL trajectories come from the actual harness rather than from a simplified proxy environment. For the Claw setting, the paper uses a 30B total-parameter MoE backbone with approximately 3B active parameters; for GUI, it uses an 8B vision-language backbone.
09
Experiments
The evaluation design deliberately spans different interaction surfaces. Claw tasks test text-based tool orchestration; OSWorld-Verified tests computer interaction; Online-Mind2Web and WebVoyager test browser interaction. This makes the framework more than a single benchmark optimization exercise. The paper is asking whether the same training infrastructure can support different harnesses, modalities, and environment requirements while retaining a common RL interface.
10
Baselines
The paper compares against both similarly sized open models and larger frontier models. For Claw, the similar-size baselines include LLaMA-4-Scout-17B-16E-Instruct, Mistral-Small-3.1-24B-Instruct, Qwen3-32B, Qwen3-30B-A3B-Thinking, and Qwen3-Coder-30B-A3B-Instruct. It also reports larger systems including Claude Opus 4.6, GPT-5.4, Gemini 3.1 Pro, Qwen3.5 397A17B, GLM 5 Turbo, MiniMax M2.7, MiniMax M2.5, and Kimi K2.5.\n\nFor GUI evaluation, comparisons include OpenCUA, UI-TARS, MolmoWeb, and other similarly sized models. The paper emphasizes that OpenForge-GUI can be competitive with or exceed models substantially larger than its 8B backbone on several benchmarks.
The baseline comparisons should be interpreted with care because benchmark protocols differ. The paper reports the exact evaluation harness and metric for each benchmark rather than collapsing them into one aggregate score. In particular, the Claw table separates pass^3, pass@3, and pass@1, while the GUI results use average success rates. This preserves the meaning of each benchmark instead of treating all percentages as interchangeable.
11
Results
OpenForge-Claw (SFT+RL) reaches 31.7 pass^3 and 55.9 pass@3 on ClawEval, 33.7 pass@1 on QwenClawBench, and 28.1 pass@1 on MCPAtlas. The SFT-only model reaches 21.7 pass^3, 52.1 pass@3, 32.1 pass@1, and 23.6 pass@1 respectively.\n\nOpenForge-GUI (SFT+RL) reaches 37.7 on OSWorld-Verified, 63.0 on Online-Mind2Web, and 72.3 on WebVoyager, compared with 34.4, 57.4, and 61.5 for SFT-only. The paper reports that these results outperform open baselines of similar size on nearly all benchmarks and, in the GUI setting, can match or surpass models several times larger.\n\nThe paper also finds that harness diversity matters. Training on ZeroClaw alone improves unseen OpenClaw and Codex performance over the base model, while training jointly on ZeroClaw, OpenClaw, and Codex produces larger gains across all three evaluation harnesses.
One of the clearest quantitative signals is the SFT-to-RL improvement. For OpenForge-Claw, SFT+RL raises ClawEval pass^3 from 21.7 to 31.7 and pass@3 from 52.1 to 55.9. It also raises QwenClawBench from 32.1 to 33.7 and MCPAtlas from 23.6 to 28.1. For OpenForge-GUI, SFT+RL improves OSWorld-Verified from 34.4 to 37.7, Online-Mind2Web from 57.4 to 63.0, and WebVoyager from 61.5 to 72.3. The consistent direction across different environments supports the paper’s claim that RL contributes useful agentic behavior rather than only task-specific imitation.
12
Ablation
The harness-transfer experiment is particularly informative because the training recipe is held constant while the harness exposure changes. A ZeroClaw-only model reaches 46.0 pass@1 on ZeroClaw, 14.7 on unseen OpenClaw, and 16.8 on unseen Codex. Multi-harness training reaches 48.5, 20.9, and 32.5 respectively. This suggests that exposure to diverse control flows can improve robustness beyond the exact harness seen during training, although the experiments do not establish that all harness diversity is beneficial.
13
Limitations
The framework introduces substantial infrastructure complexity: remote rollout containers, Kubernetes orchestration, inference proxies, networking, environment provisioning, and failure handling all become part of the training system. The paper also notes that rollout failures can arise from network issues, harness crashes, or timeouts rather than from the policy itself. Its current strategy discards trajectories that terminate with such errors, leaving better credit assignment for partial rollouts as future work.\n\nThe experiments are also limited to the selected harnesses, environments, benchmarks, and relatively small task collections. The reported gains demonstrate the value of harness-native training, but they do not establish that every production harness or every agent domain will benefit equally. Finally, error recovery remains weak even after RL, showing that improving long-horizon reliability is not solved by the framework alone.
There is also an important distinction between benchmark reliability and deployment reliability. Remote execution makes large-scale training possible, but the system itself becomes a distributed application with multiple independent failure sources. The paper’s decision to discard infrastructure-failed trajectories is pragmatic, but it can reduce data efficiency. Production systems would likely need richer failure classification, replay, provenance, and observability so that environment failures are not confused with policy failures.
14
Conclusion
OpenForgeRL tackles a practical problem in agent development: the system used to train an agent is often much simpler than the system used to deploy it. By keeping the real inference harness and environment intact, then connecting them to standard RL infrastructure through a proxy and remote rollout orchestration, the framework reduces this mismatch.\n\nThe empirical results show meaningful improvements from SFT+RL and from training across multiple harnesses, while behavioral analysis suggests that RL can improve self-verification, tool coverage, and multi-step reliability. At the same time, error recovery remains weak and the infrastructure itself introduces operational complexity.\n\nFor production AI engineers, the main lesson is architectural: the harness is part of the agent, not merely glue around the model. If tools, control flow, context management, and environment interfaces materially shape behavior, those components need to be represented in the training and evaluation loop.
The broader engineering lesson is that agent training infrastructure increasingly resembles distributed systems engineering. Model optimization is only one part of the loop; environment lifecycle, tool interfaces, reward verification, rollout scheduling, state capture, and failure recovery can all affect the learned policy. OpenForgeRL provides an interface for bringing those concerns into the training loop without forcing every new environment to become a custom RL implementation.
15
References
Primary source: Yu et al., “OpenForgeRL: Train Harness-native Agents in Any Environment,” arXiv:2607.21557v3, revised August 7, 2026.\n\nKey external and related sources include Orchard, Polar, SkyRL-v0, OpenWebRL, SWE-RL, Search-R1, OSWorld, WebVoyager, MCPAtlas, ClawEval, QwenClawBench, veRL, and OpenAI/Anthropic documentation cited by the paper.\n\nThe paper reports that code, data, and models are released through the project resources.
Source trail
Primary papers, implementation pages, and external source records used to ground this explainer.
OpenForgeRL: Train Harness-native Agents in Any Environment
Yu et al., arXiv:2607.21557v3 (2026)
Primary paper.
Orchard: An open-source agentic modeling framework
Peng et al., 2026
Underlying environment and orchestration framework referenced by OpenForgeRL.
Polar: Agentic RL on Any Harness at Scale
Xu et al., 2026
Concurrent harness-level RL work discussed by the paper.
SkyRL-v0: Train Real-world Long-horizon Agents via Reinforcement Learning
Cao et al., 2025
Long-horizon agent RL infrastructure cited by OpenForgeRL.
01arXiv paper02OpenForgeRL GitHub03OpenForgeRL project04Microsoft Research publication page05SkyRL06PolarContinue reading
Related research
Agentic Reasoning
A tool-using agent framework that extends LLM reasoning with web search, coding, and structured reasoning memory for deep research tasks.
AI AgentsGenerative Agents
An agent architecture combining memory, retrieval, reflection, and planning to simulate believable long-horizon behavior.
AI AgentsMIRA
MIRA is an autonomous medical agent evaluated in a sandboxed EHR workflow with tools for diagnosis, testing, treatment, and admission decisions.