MachinoAI explainer / AI Agents
Agent-Editing World Model: Rethinking World Modeling for LLM Agents
AEWM reframes world modeling for LLM agents from predicting tool observations to judging and editing the agent state that drives future decisions.
01
Abstract
AEWM proposes a different target for language world models. Instead of learning to reconstruct future tool observations, it learns how an agent’s current reasoning-action continuation is likely to affect future task progress. The motivation is that search results, terminal output, and test results are high-entropy and execution-dependent; when the real environment is available, fabricating those observations can be less useful than correcting the interpretation that produced the next action. AEWM has two capabilities: Action Judge labels a proposed continuation as CRITICAL, EXPLORATORY, or NOISY; State Revision rewrites a NOISY continuation using the same visible history. EditAct inserts this intervention before real execution, and AEWM-RFT uses verified EditAct trajectories to train an agent to internalize the correction behavior. On a 3,000-decision Action Judge benchmark, AEWM reaches 70.5% macro-F1, 10.6 points above the strongest compared frontier baseline. EditAct improves average task scores by 3.2–6.7 points across six benchmarks and three Qwen3.5 backbones.
02
Introduction
Long-horizon agents are not only prediction systems; they are stateful decision processes. Every reasoning trace and tool action becomes part of the history used for later decisions. A wrong hypothesis can therefore become persistent state: an agent may treat an unverified lead as fact, keep an obsolete plan after contradictory evidence, or mistake partial progress for completion. Real observations do not automatically repair these mistakes because the agent still has to interpret them. The paper calls this failure mode task-state contamination. Existing language world models generally borrow an environment-centered objective: predict what a tool or environment will return after an action. AEWM asks a more agent-centered question: given what the agent currently knows and what it is about to do, will that decision help, explore productively, or reinforce a bad direction? This reframing makes the world model an intervention mechanism rather than a simulator.
03
Problem
The core problem is persistent decision error under partial observability. Let the visible interaction history be h_t and the agent propose reasoning r-hat_t and action a-hat_t. The proposal is not merely an ephemeral thought: once executed and observed, it becomes part of h_{t+1}. If the proposal embeds an unsupported assumption, subsequent steps inherit it. The system therefore needs a way to judge a proposal before execution and, when necessary, replace both the reasoning and action rather than merely attaching a critique. The paper also argues that observation prediction is a poor fit for many tool environments because search rankings, web content, filesystems, runtime state, and test outcomes are difficult to predict exactly and are already available from real execution.
04
Background
Paper-derived background: conventional world models learn action-conditioned environment dynamics; language world models extend this idea to tool or web observations. AEWM instead connects to process supervision, revision learning, ReAct-style tool use, and world models for web/code agents. The paper distinguishes three decision roles. CRITICAL decisions close an important gap or perform a necessary state change. EXPLORATORY decisions reduce uncertainty or test a plausible branch and therefore should not be suppressed merely because they are not on the shortest path. NOISY decisions have little expected progress, repeat unproductive behavior, violate constraints, or reinforce an incorrect direction. External research: the surrounding literature includes observation-predictive language world models, web agents with world models, Reflexion-style verbal feedback, and process verifiers. AEWM differs by making downstream decision effect and state editing the central modeling target.
05
Methodology
The method has three linked stages. First, Action Judge data are synthesized from successful verified trajectories. An annotation agent labels individual turns using the corresponding observation and later task outcome, while filtering checks label consistency, action-observation alignment, action validity, and grounding in information visible before execution. Second, State Revision data are generated by proposing a task step, using an Action Judge checkpoint to find NOISY proposals, generating a replacement from the same history, executing it in the real environment, and retaining it only when the new step produces substantive progress. Third, EditAct runs online: the base agent proposes a reasoning-action pair; AEWM judges it; CRITICAL and EXPLORATORY proposals are retained, while NOISY proposals are replaced by State Revision; the selected action is then executed in the real environment.
Formula notes
Mathematical details
Conventional observation prediction vs AEWM state editing
A conventional language world model predicts the next observation; AEWM instead transforms the current pre-execution state into an edited state containing revised reasoning and action while preserving the observed history.
EditAct decision rule
Keep productive or useful exploratory proposals; replace only proposals judged noisy before executing the selected action in the real environment.
Grounded history update
The selected action is executed by the real environment, and its actual observation becomes the evidence used for the next decision.
06
Architecture
The AEWM architecture is an intervention loop around an ordinary agent. The base agent proposes (r-hat_t, a-hat_t) from the current history h_t. AEWM Action Judge reads the pre-execution state and emits one of three decision-effect labels. If the label is CRITICAL or EXPLORATORY, the proposal proceeds unchanged. If it is NOISY, State Revision generates (r-tilde_t, a-tilde_t) from the same visible history. Only the selected action is executed. The real environment returns observation o_t, which is appended to the history before the next agent turn. This matters because the model edits the state that future reasoning will consume rather than merely producing a side-channel score. Figure 2 in the paper shows the training-data synthesis, AEWM training, EditAct inference, and AEWM-RFT loop.
07
Dataset
The paper uses three domains: Search, Terminal, and Software Engineering. Search training uses internal deep-search data; Terminal uses CalibForge; SWE uses DeNovoSWE. Action Judge data are generated with DeepSeek-V4-Pro and GLM-5, while State Revision uses Qwen3.5-35B-A3B and DeepSeek-V4-Pro as proposal/revision agents. The held-out Action Judge benchmark contains 3,000 decisions, equally split across Search, Terminal, and SWE, and is built from verified trajectories using repeated annotation, consistency filtering, model-based review, and diversity-aware sampling. Downstream evaluation uses BrowseComp, DeepSearchQA, Terminal-Bench 2.0, Doc2Repo, NL2Repo, and SWE-Bench Pro. DeepSearchQA and SWE-Bench Pro are identified as out-of-distribution relative to the training mix.
08
Training
AEWM uses a two-stage training recipe. Mid-training uses approximately 52B tokens across the three domains, mixing original agent trajectories with synthesized Action Judge and State Revision data. The paper reports 20.56B tokens for Action Judge supervision, 25.01B for State Revision supervision, and 6.59B raw trajectories in the overview. SFT then uses 120K curated examples: 60K Action Judge and 60K State Revision, with 40K examples from each domain. The AEWM backbone is Qwen3.5-35B-A3B. A separate transfer stage, AEWM-RFT, retains verified high-quality EditAct trajectories and fine-tunes the agent on those trajectories. The purpose is to internalize useful state-correction patterns so the deployed agent can run ordinary ReAct without an online AEWM call.
09
Experiments
The authors run three independent evaluation runs per downstream benchmark and report mean scores. EditAct is implemented in ReAct-based Search, CalibForge-Eval Terminal, and SearchSWE software-engineering scaffolds. The three inference agents are Qwen3.5-4B, Qwen3.5-9B, and Qwen3.5-35B-A3B. Action Judge is compared with Gemini-3-Pro, GLM-5.2, Qwen3.7-Max, GPT-5.5, and DeepSeek-V4-Pro using accuracy and macro-F1. EditAct is compared with ReAct, step-level Best@3, and trajectory-level Best@3. Within each comparison, methods share the same agent, tasks, environment settings, and evaluation criteria.
10
Baselines
The principal downstream baseline is ReAct, which reasons and acts against real tools but does not selectively edit its current reasoning-action continuation. Step-level Best@3 samples three candidates at each decision and uses a verifier to select one; trajectory-level Best@3 generates up to three complete trajectories and submits the verifier-selected result. These baselines test whether AEWM is merely benefiting from extra candidate generation. The transfer baseline Self-RFT trains Qwen3.5-35B-A3B on its own verified trajectories, whereas AEWM-RFT learns from trajectories improved by EditAct. Action Judge is separately compared with strong frontier models.
11
Results
AEWM reaches 70.5% overall macro-F1 on Action Judge, versus 59.9% for DeepSeek-V4-Pro, a 10.6-point gap. Domain macro-F1 is 60.9% Search, 72.1% Terminal, and 77.8% SWE. On six downstream benchmarks, EditAct averages 41.8 for Qwen3.5-4B, 44.1 for Qwen3.5-9B, and 48.8 for Qwen3.5-35B-A3B; strongest competing averages are 35.1, 38.9, and 45.6, giving gains of 6.7, 5.2, and 3.2 points. The full Qwen3.5-35B-A3B EditAct row is 48.1 BrowseComp, 78.6 DeepSearchQA, 48.3 Terminal-Bench 2.0, 48.9 Doc2Repo, 25.3 NL2Repo, and 43.8 SWE-Bench Pro. AEWM-RFT scores 45.4/43.4/48.6 on BrowseComp/Terminal-Bench 2.0/Doc2Repo, versus 43.2/40.8/46.1 for Self-RFT.
12
Ablation
Table 3 isolates the source of the gains. On BrowseComp, Terminal-Bench 2.0, and Doc2Repo, full AEWM scores 48.1, 48.3, and 48.9. Random gating is lower at 44.1, 40.8, and 43.9, supporting learned intervention selection. Agent resampling and AEWM Hint are also below the full system, supporting direct state editing over generic regeneration or guidance. Action-only and reasoning-only revision do not match joint revision. Self-WM and DeepSeek-V4-Pro WM are below AEWM on the three reported tasks. Training ablations show that full mid-training plus SFT beats SFT-only by 0.8, 6.4, and 4.3 points, and beats mid-training-only by 0.3, 4.1, and 1.9 points on the same benchmarks. Figure 5 adds a behavioral signal: the longest search-only sequence falls from 16.9 to 13.2 while distinct requested pages rise from 7.7 to 10.1 with AEWM.
13
Limitations
The paper does not provide a standalone limitations section, so limitations are separated into paper evidence and engineering interpretation. Paper evidence: Qwen3.5-Plus gains are strong on BrowseComp but limited on Terminal and SWE, which the authors attribute to a capacity gap between the stronger agent and AEWM. A learned gate can also make false interventions: suppressing a useful action is costly, while missing a noisy action leaves contamination unresolved. Evaluation covers Search, Terminal, and SWE rather than GUI, robotics, multimodal, or multi-agent settings. Engineering interpretation: EditAct adds online model calls, so latency, memory, token cost, and recovery overhead should be measured before production deployment; the paper does not provide a complete wall-clock or cost analysis for this extra loop. The 3,000-decision judgment benchmark is carefully constructed but depends on trajectory-derived labels and annotation/filtering quality.
14
Conclusion
AEWM reframes world modeling for LLM agents around the state that controls future decisions. Its strongest idea is not simply adding a critic; it is to distinguish useful exploration from harmful continuation, revise reasoning and action together when necessary, and then obtain the next observation from the real environment. Results support this design across Search, Terminal, and SWE, with consistent gains across three agent scales and transfer through AEWM-RFT. The production lesson is to treat task state as an explicit reliability surface. If stale assumptions, failed plans, or partial completion can persist in history, an intervention layer should be evaluated on whether it improves the committed trajectory, not only whether it produces a plausible critique.
15
References
Primary: Shuang Sun et al., Agent-Editing World Model: Rethinking World Modeling for LLM Agents, arXiv:2609.28416, 2026. Related references used by the paper include ReAct, Reflexion, General Agents Need World Models, Web Agents with World Models, ECHO: Terminal Agents Learn World Models for Free, SWE-Master, CalibForge, and the cited Search/Terminal/SWE benchmarks. External research checked during this run includes the official arXiv paper record and the public RUC AI Box repository listing for the implementation.
Primary links and external source pages used by this explainer.
Agent-Editing World Model: Rethinking World Modeling for LLM Agents
arXiv:2609.28416 (2026)
Primary paper.
Agent-Editing World Model GitHub
RUCAIBox/Agent-Editing-World-Model
Public implementation repository identified during external research.
ReAct: Synergizing Reasoning and Acting in Language Models
ReAct
Core agent baseline cited by the paper.
Reflexion: Language Agents with Verbal Reinforcement Learning
Reflexion
Related work on language-agent self-improvement cited by the paper.
arXiv paperFull paper PDFOfficial GitHub repositoryContinue reading
Related research
Generative Agents
An agent architecture combining memory, retrieval, reflection, and planning to simulate believable long-horizon behavior.
AI AgentsMIRA
MIRA is an autonomous medical agent evaluated in a sandboxed EHR workflow with tools for diagnosis, testing, treatment, and admission decisions.
AI AgentsAgentic Reasoning
A tool-using agent framework that extends LLM reasoning with web search, coding, and structured reasoning memory for deep research tasks.