
Stale assumptions in the history bend later decisions
Hand an agent a task with many steps and the shape of its failures changes. It does not simply get one action wrong. An assumption it made partway through goes stale, stays in the history, and quietly bends every decision that follows. The authors call this task-state contamination. Unsupported beliefs, plans whose conditions have since changed, guesses never checked — in the history, all of them sit alongside verified facts and look the same.
The paper's second observation concerns how language world models have been built. Most have tried to predict the observation a tool returns. But much of that response is only settled by executing it. Where real execution is available, effort spent predicting it earns little. Starting there, the authors change what is predicted[1]. The primary source for this article is an arXiv preprint that has not been peer reviewed.
Modelling how progress changes, not what the tool says
The core of the proposal is one substitution: instead of simulating tool responses, model how reasoning and actions change the state of progress on the task. The authors call the framework the Agent-Editing World Model, or AEWM.
AEWM combines two parts. Action Judge sorts the decision at hand into critical, exploratory, or noisy. State Revision then rewrites the reasoning-and-action continuation that was marked noisy, working from the same observed history. It does not attach a critique; it replaces the continuation itself.
EditAct is the system that puts those two together with real execution. Rather than handing back a critique for the agent to read, it changes the state on which subsequent decisions rest. That is the break with the advisory pattern. Training is described as spanning three domains — Search, Terminal, and Software Engineering — through mid-training and supervised fine-tuning.
What the abstract claims
The abstract reports the following. On the authors' own Action Judge benchmark, AEWM reaches 70.5% macro-F1, exceeding the strongest frontier baseline by 10.6 points. Across six benchmarks and three agent backbones, EditAct is said to improve average scores by 3.2 to 6.7 points over the strongest baseline[1].
The authors additionally describe rejection sampling fine-tuning on verified EditAct trajectories, which they call AEWM-RFT, as improving over Self-RFT by 2.2 to 2.6 points across three domains — and doing so without online AEWM guidance. The claim, in other words, is that once the machinery has been consumed as training material it can be removed. All of this is what the authors assert; none of it is an established result.
What it looks like in practice
Editing a history sounds abstract. In practice it comes down to this difference.
The advisory pattern
System:"That assumption may no longer hold."
Agent:(the original assumption stays in the history; the warning sits next to it)
Agent:(on the next decision, rereads the original assumption and continues the same way)
The state-editing pattern
System:(rewrites the continuation it judged noisy, from the same observed history)
Agent:(the next decision is taken on the rewritten state)
Agent:(there is no occasion to reread the original assumption)
What matters here is not accuracy but a design fork: do you delete the error, or keep it and annotate it? Deleting is fast, but the record of the deleted judgement thins out. Annotating preserves the record, and the record keeps pulling.
What is not new, and what cannot be read
Mechanisms in which a model revisits and repairs its own output have been accumulating for some time[4], and interleaving reasoning with action is already standard construction[5]. What this paper adds is passing the result of that revisiting back as a replaced history rather than as advice, and narrowing what gets rewritten by sorting decisions into three kinds.
Plenty cannot be read from the abstract. How the boundary for calling a decision noisy was set; what happens when a rewrite rests on a mistaken judgement; what becomes of the history's coherence as rewrites accumulate. The abstract states no conditions for any of these. That the benchmark is the authors' own is a further reason to wait for independent replication. Because this is a preprint, judgement belongs after the full text and review.
Votes arrived; a cluster did not
The paper drew 19 upvotes and 2 comments, and an implementation is public with 6 stars[2][3]. Some readers judged it worth reading and a few people have picked it up.
No other paper on the same question appeared in the same window within the range of this collection. The cluster size is 1. This site measures the spread of a topic across independent signals: reader votes, implementations, press, and a cluster of papers. Votes are neither correctness nor importance. They mean that several people thought it worth reading, and nothing beyond that. One further caution: now that "world model" is a widely used phrase, papers can agree on the words while differing in substance.
Where this touches promotional material review
Reviewing promotional material is a long-horizon task. A finding is raised, a revision comes back, the revision is checked against the rest of the document, and it goes round again. Assumptions made partway through go stale in that loop more often than not. The package insert was revised; the conditions attached to a cited figure changed; the basis for wording cleared in an earlier version moved to a different document. Unnoticed, each of these keeps acting on the next decision in its stale form.
That failure mode — the stale assumption left in the history — is what this paper addresses. But the remedy cannot be carried into a regulated setting as it stands. A review record has to retain mistaken judgements as well as correct ones. If what was changed and why cannot be traced, nothing can be explained to an auditor, and a repeat of the same mistake cannot be detected.
So the part that transfers is not the rewriting. It is the sorting of decisions into kinds. Which finding is critical, which is exploratory, which is noise that can be dropped. As long as people make that distinction, the criteria live only in the reviewer's head. Writing those criteria down has value independent of whether any machine is introduced. And then, having written them down, choose annotation over rewriting. That is the order.