A terminal agent generates and runs shell commands, and a poor action stays in the environment, unlike a chat reply that can be corrected. Mid-Harness samples several candidate actions between the model's generation and the harness's execution, and runs only the one a verifier selects. The gain depends on the verifier's strength rather than the number of candidates; under weak verification more candidates help little. With GPT-5.6 Sol as verifier and 8 sampled actions, Pass@1 rose from 50.00% to 68.03%. This is a preprint result, to be treated as design input.
Image abstract — the whole article on one page (click to enlarge)

1. One bad step shapes everything after it

A terminal agent generates a shell command, executes it, reads the output and decides on the next command. Because model generation is stochastic, the same situation yields slightly different commands each time.

The authors start from the observation that the ability to generate a useful action does not ensure that the action is executed reliably. Their example is installing the wrong package. A poor command rewrites the environment itself, so later progress is hindered even if the model could have produced a better alternative. In a chat, a bad reply can be corrected in the next message; in a terminal, the operation leaves a mark on the environment.[1]

This leads them to ask where test-time compute should be spent. The usual choices are letting the model think longer internally, or rerunning the whole task several times and keeping the best result. This paper examines a third option in between: reselecting at the level of each single action. The primary source is a preprint posted on arXiv and has not yet been through peer review.

2. The core idea: verify at the boundary, then execute

The core of Mid-Harness fits in one sentence: between the moment the model generates an action and the moment the harness executes it, sample several candidate actions, and forward only the one a verifier selects.

Neither the model that generates actions (the generator) nor the system that runs commands (the harness) is modified. Mid-Harness is a layer inserted between them, so it can be tried without changing the structure of an existing agent.

Several verification mechanisms are compared. The abstract contrasts using a stronger external model as the verifier with using the generator model itself as the verifier. In the latter case, pairwise verification, which compares candidates two at a time, performs best among the mechanisms evaluated. In addition, distilling the stronger verifier's responses into the same small model improves results further while the action generator stays unchanged.

3. What the abstract reports

With TMAX-9B as the generator, sampling more actions gives little benefit when verification is weak. With a capable verifier, by contrast, the agent can exploit useful alternatives that the same generator had already produced, according to the authors.

One concrete figure is given. On TerminalBench-Lite, using GPT-5.6 Sol as the verifier with 8 sampled actions raises Pass@1 from 50.00% for the base agent to 68.03%. Pass@1 is the share of tasks solved in a single attempt.

A second claim concerns cost. With TMAX-9B, combining action scaling with trajectory scaling reaches higher success at a lower estimated token cost than generating more trajectories alone. Mid-Harness is also said to improve results across additional models, benchmarks and harnesses, but the abstract does not name them or give figures.

4. A worked example: setting up an analysis environment

Consider an agent asked to set up a fresh environment for statistical analysis, install the required libraries, and get an existing analysis script running.

Base agent

Agent: Installing the required library. (It generates a command for a different package with a similar name and runs it immediately.)

Agent: The script fails. Fixing dependencies. (The package it installed has changed the version of another library, and each fix breaks something else.)

With Mid-Harness

Agent: Generating several candidate install commands.

Verifier: One candidate points to a similarly named but different package. Forwarding the candidate that names the intended library.

Agent: Executed the selected command. Moving on to run the script.

Note that the generator's ability has not changed. The correct command could appear among the base agent's candidates too. The difference is whether the decision about which candidate touches the environment is made before execution. And, as the abstract's results show, if that decision is made poorly, more candidates do not help.

5. What is not new, and what the abstract does not settle

Generating many candidates and letting a verifier choose is not a new idea. For math word problems it was shown earlier that sampling many solutions and selecting with a trained verifier improves accuracy.[4] The basic pattern of agents that interleave reasoning and acting was established by work such as ReAct.[5] The contribution here can be read as moving that selection from the level of whole tasks to the level of a single action, placing it at the model-harness boundary, and studying what makes it work.

There are limits. First, the abstract does not describe the number or composition of tasks in TerminalBench-Lite. The name resembles Terminal-Bench, an evaluation suite for terminal tasks,[3] but the relationship cannot be confirmed from the abstract. Second, the largest gain comes from using a strong external model as the verifier, and the abstract does not make clear how much of the verifier's own cost is included in the estimated token cost comparison. Third, nothing is said about what happens when the verifier picks a wrong candidate, for instance how often an irreversible operation slips through. Fourth, no code repository was registered at the time of collection.

On the paper-sharing page the paper has 111 upvotes and 5 comments, and 0 GitHub stars.[2] Upvotes show how widely it was read, not whether its claims are right. No independent reproduction has been found.

6. Why this is getting attention, as a paper without a cluster

In this collection run, the paper's cluster size was 1, and only one signal family, reader votes, was raised. In other words, this is not a case of several independent papers taking up the same question at once; it is a single paper that attracted votes. It should be read as a pick made under that condition.

Even so, its keywords (actions, harness, scaling, terminal) reflect a recent direction. This site has recently covered a run of preprints on designing and self-improving the harness that wraps an agent. Where earlier work offered two places to spend inference compute, inside the model or across repeated whole attempts, this paper adds a third: the boundary with the harness.

The other two papers selected on the same day also deal with what an agent relies on when choosing its next move in the middle of a long task. They are separate papers that appeared at the same time, however, not a cluster that cites one another.

7. Where this connects to pharmaceutical and regulatory work

The first connection is the idea of checking before acting. In pharmaceutical operations, confirming a change against an approved procedure before applying it to a system is standard practice. Mid-Harness inserts something similar before each of the agent's actions. It does not replace human review, but it is useful input when deciding, for agents placed in business environments, before which operations a verification step should sit.

The second is the verifier's record. If the system keeps which candidates were generated, which was chosen and why the others were rejected, the course of operations can be traced later. The abstract does not describe what is logged, so users would need to build that themselves.

The third is dependence on verifier quality. Within this paper, adding candidates under a weak verifier gave little benefit. Having a verification step is not in itself grounds for safety; the verifier needs its own evaluation. Given that this is a preprint, the result is better treated as input for design than as grounds for deploying agents in work that includes irreversible operations.