1. Assisted successes are hard to use as training data as they are

Terminal agents type commands, manipulate files, and build software, and they fail more often as tasks get harder. The most direct way to improve them on hard tasks is to collect records (trajectories) of the tasks actually being solved and train on those.

The catch is that many successes on hard tasks rely on a specialized harness. A harness is the setup placed around the model: task-specific guidance, extra tools, interventions along the way. The authors point out that these specialized interventions may not be available when the model is actually deployed. Training on successes that depended on assistance does not guarantee the same behavior in a deployment without it.[1]

The paper asks how to keep the value of those successful records while removing the dependence on assistance. The primary source is an arXiv preprint and has not yet gone through peer review.

2. The core idea: turn successes into runbooks and redo them in a general setup

RSR comes down to one move: rebuild successes found under specialized harnesses as trajectories recorded under a general harness. A single base model, Qwen-3.8-27B, is used throughout, from discovery to rewriting.

The rebuilding involves three roles. A planner extracts the procedure from a successful trajectory into a runbook. A critic checks the runbook for leakage of verifier details or of the solution itself, and when it finds a problem it has the runbook revised; repeating this is what "recursive" refers to. An executor then follows a runbook that has passed the critic, in a fresh sandbox under the general harness.

The critic matters because if answers or grading logic slip into a runbook, the re-execution only looks successful. Routing everything through a runbook keeps the benefit of the assistance as a written procedure while removing the assistance itself from the conditions the model trains under.

3. What the abstract reports

The authors worked with approximately 3K terminal tasks they curated themselves. Three harnesses together solved 759 tasks, 34.3% more than the strongest individual harness in the recorded pool. Using several harnesses widened the set of successes on its own.

RSR expanded 2,001 successful source trajectories into 11,094 rewritten trajectories for supervised fine-tuning. A model trained on these trajectories is reported to outperform both the base model and a model fine-tuned directly on the source trajectories.

Evaluation uses the Terminal-Bench family of terminal-task benchmarks.[4] Compared with the base model, pass@3 rose from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on the authors' own Terminal-Bench Hard, and from 3.0% to 6.0% on their own Software Terminal-Bench. Process reward on Long-Horizon Terminal-Bench rose from 0.21 to 0.29. The abstract does not give the size of the gap between RSR and direct trajectory fine-tuning.

4. A worked example: an agent for internal data-platform work

Consider a team that wants an agent to handle fixed transformation and verification jobs on an internal data platform.

Developer: In the evaluation environment it solves quite hard jobs. With the special guidance and extra tools switched on, it does well.

Quality assurance: Will that guidance and those tools be in production too?

Developer: No. Production has only the standard tools. That is why we want to train on the successful runs, so it can solve the jobs with standard tools alone.

Quality assurance: Have you checked that those runs do not contain the evaluation answers or the verification logic? If they do, the model just learns to solve while looking at the answer. And please do not evaluate the trained model only on jobs from the same source you trained on.

The conversation maps onto RSR's three roles: extract the success as a procedure, check for leaked answers or verifier details, and redo the work in the standard environment. The last remark also anticipates a limit discussed in the next section.

5. What is not new, and what the abstract leaves open

Retraining a model on its own successful outputs is not a new idea. In reasoning research, collecting a model's own rationales that reached correct answers and training on them has been known for some time.[5] Selecting successful candidates from many samples to build supervised data is also common practice. The contribution here is the focus on the gap between assisted success and deployment conditions, and a pipeline that rewrites through runbooks and screens for leakage.

There are several limits. First, the training tasks were curated by the authors, and two of the four evaluation benchmarks are also their own. When training and evaluation tasks come from the same source, gains can look larger than they really are. Second, leakage screening is done by the critic model, and the abstract does not say how often it misses something. Third, on Terminal-Bench 4 and Software Terminal-Bench, where the starting success rate is low, the improved rate is still low; this is not yet a level at which one could say the model can do those hard tasks. Fourth, the comparison concerns mainly one base model, and it is unknown whether other models would show the same effect. A public implementation could not be confirmed from the abstract either.

6. Why this topic is clustering now

The paper drew 74 upvotes on Hugging Face Daily Papers.[2] In this site's selection, two signals fired, reader votes and clustering, and 2 papers in the same category formed one cluster. Upvotes are votes from readers who thought a paper worth reading; they do not establish that it is correct or important.

The other paper in the cluster is VeriHarness. It addresses how to verify the output of agents doing long tasks without reference answers or grading rubrics, by giving the same LLM used for generation a workspace and evidence-gathering tools so it can act as a verifier.[3] RSR works on the side of producing training data and VeriHarness on the side of checking outputs, but both start from the same question: how to raise the quality of long tasks by changing the setup around a fixed base model. This site has recently covered several preprints on terminal agents and harness design, so the question is clearly being raised by more than one group. Still, the cluster size is 2, which is not enough to call it a thick topic.

7. What connects to pharmaceutical and regulatory work

Pharmaceutical companies are also looking at handing routine data processing and document work to agents. The issue this paper addresses, a mismatch between conditions at evaluation and conditions in production, is one that validation of such systems will always raise.

Two points carry over. The first is to check that the environment in which the agent succeeded during evaluation matches the production environment. If special guidance or tools were present only during evaluation, the scores do not represent production performance. The second is to screen training data for leaked answers or verification logic, and keep a record of that screening. Keeping the provenance of training data and the results of the screening makes it easier to explain the model's behavior later if that becomes necessary. RSR's design of rewriting through runbooks also fits operations based on standard operating procedures, because it leaves the procedure behind as a document.

This paper is a preprint, and the practical implications above go only as far as its own scope supports.