
1. The difficulty of inferring how an unfamiliar world works
An agent in an unknown environment has to work out what its actions do from its own observations. The authors begin with a specific problem: when observations are limited, several explanations of the world can fit everything seen so far yet predict different outcomes in states not yet visited.[1]
Language models carry a lot of general knowledge, but they do not know the particular rules of the environment in front of them. Simply remembering past episodes does not help much when a new situation looks different. Committing early to one explanation has the opposite risk: the agent keeps acting on a wrong rule.
The question the paper takes on is whether an agent can write down the rules of its environment explicitly and keep revising them against evidence, without updating the model's weights.
2. The core idea: write a rulebook, compile it, and verify by replay
Memento 3 builds on earlier work in the same series and lets a frozen LLM agent learn an explicit world model through external memory. At the center is a natural-language rulebook the agent keeps as persistent memory. It records revisable hypotheses about how the environment behaves and leaves unknown parts deliberately underspecified.
The agent compiles the rulebook into executable code and uses that code to predict the next state and plan actions. It runs a continual loop of observation, reflection, rule revision, compilation, and verification, using prediction errors to refine both the rulebook and the code.
The acceptance rule matters most. Updated code is accepted only when the language model judges it faithful to the rulebook and a cell-exact replay reproduces every observed state transition. The paper also describes a population extension that keeps multiple world models in parallel, shares interaction evidence among them, and uses disagreements between their predictions to guide exploration.
3. What the abstract reports
The main evaluation uses ARC-AGI-3, an interactive reasoning benchmark in which agents have to explore environments they have not seen before.[4] According to the abstract, the single-model agent clears every level of all 25 public games, reaches a mean Relative Human Action Efficiency (RHAE) of 100.0, and uses 44% of the human action count.
A second case study uses Atari Pong. A feedback controller learned by the agent wins 21:0 in each of three evaluated episodes with different openings, without any further LLM calls during play.
The authors frame these results as a process in which the agent explores on its own, revises its world model, and uses verified updates to guide the next round of interaction and learning, while the underlying language model stays fixed.
4. A concrete case: learning the rules of a new data entry system
Suppose an agent is assigned data entry in a new electronic system for clinical trial data. The manual describes only some of the conditions that trigger a validation warning.
An agent that stores raw experience
It accumulates examples of entries that produced warnings and applies them to similar inputs. Because nothing records why a warning appeared, its predictions fail on inputs that differ slightly. When a prediction fails, there is no written belief to point to and correct.
An agent that keeps a rulebook
It writes hypotheses such as "a warning appears when one field is blank and a related field is filled," and leaves unknown conditions open. It turns the rulebook into checking code and adopts a revision only if the code reproduces every past entry and warning. A failed prediction stays on record as a counterexample, which shows which rule needs fixing.
With the second approach, what the agent believes can be read as text. A person can look at it and say "this rule contradicts the manual." The first approach gives no such opening.
5. What is not new, and what the abstract does not settle
Learning a model of the environment and using it for planning has a long history in model-based reinforcement learning. Writing world rules as programs has also been tried before, and the paper itself states that it extends the Memento series. What appears new is the pairing of a natural-language rulebook with executable code, the two-part acceptance test that combines an LLM judgment with exact replay of past observations, and the framing of this loop as one route to recursive self-improvement.
Limits are visible. First, the ARC-AGI-3 results concern the public games. The abstract says nothing about held-out evaluation games. Second, the abstract does not define RHAE or explain how a mean of 100.0 should be read. Third, cell-exact replay works because the environments have discrete, deterministic states. How the check would be replaced in noisy or stochastic settings is not described.
Fourth, the abstract reports no compute cost, such as the number of LLM calls or wall-clock time. Fifth, the Pong case is one game and three episodes. Finally, the judgment of whether code is faithful to the rulebook is made by the language model itself, and the abstract offers no measure of how reliable that judgment is.
6. Why this topic is forming a cluster now
In this collection run, the paper shares a cluster with Agent Plasticity, for a cluster size of 2.[3] That paper studies agents that turn past experience into reusable artifacts and defines plasticity as how efficiently an agent converts experience into gains on held-out interactions. Its authors report that models with similar opportunities to learn follow very different improvement paths, and that gains often transfer only partly outside the training regime.
Both papers point in the same direction: improving agents through external memory or artifacts rather than weight updates. Memento 3 offers a mechanism; Agent Plasticity offers a way to measure such mechanisms. Read together, they make clear that the open question for Memento 3 is how far its learning carries beyond the environments it was trained in.
On the paper-sharing site, Memento 3 has 22 upvotes and 3 comments.[2] No public code repository was listed at collection time. Upvotes reflect researchers' attention; they do not show that the claims are correct.
7. Where this connects to pharma and regulatory work
For anyone running systems under regulation, a learning agent is hard to manage because its behavior can change without anyone noticing. When a model's weights are updated, it is hard for a person to read what changed.
In the Memento 3 design, what changes is the rulebook and its code, not the model. The rulebook is readable text, and each update carries evidence that it reproduces past observations. In principle, this makes it possible to treat "rulebook and code versions" as the object of change control and to record which observation led to which rule revision.
Because the acceptance test is replay of past observations, though, nothing guarantees correctness in situations not yet observed. This mirrors a familiar distinction in validation between fitting historical data and working correctly in future use. Given that this is a preprint, any organization considering such agents should decide beforehand who approves rulebook revisions and at what point self-improvement stops for human review.