
1. Search agents make the same small decisions over and over
An agent that looks things up does more than search and write an answer. Along the way it keeps making short decisions: whether a retrieved document is relevant to the question, whether there is enough evidence to answer, and what to search for next. The authors focus on these.[1]
In many agents, these decisions are handed to a generative language model too. The model writes out its reasoning and conclusion token by token, and the agent reads that text to decide what to do next. The abstract names two problems with this. Generating text for every short decision adds latency that piles up, and the confidence the model expresses is not reliable.
The second problem is easy to overlook. If a model says "this document is relevant" and there is no trustworthy sense of how sure it is, there is no principled way to decide when to pass the judgment to a person or a larger model.
2. The core idea: score the options instead of generating text
SearchJev is a fast model that handles decisions only. The authors call it a System-1 model, borrowing the familiar split between fast and slow thinking, and separate it from System 2, which keeps planning, query generation, and answer composition.
The mechanism is straightforward. Given the current search state and a decision schema that defines the allowed options, SearchJev assigns a score to each legal option directly. It skips token-by-token generation, which is where the speed comes from.
For confidence, the authors propose Soft-Label Learning for Calibrated Decisions (SLCD). It learns option probabilities from supervision whose correct answer is itself uncertain, and calibrates those probabilities so they track how often the decision is actually right. In the dual-system agent, SearchJev handles the short decisions and delegates only uncertain judgments to System 2, which keeps the planning and writing work.
3. What the abstract reports
For training and evaluation, the authors introduce SearchDecision-Bench, which brings six types of search decisions into a single benchmark.
According to the abstract, on this benchmark SearchJev produces better decisions than same-size Qwen3.5 autoregressive models, makes decisions 5.2 to 5.3 times faster, and lowers average expected calibration error (ECE), a measure of the gap between stated confidence and actual accuracy, by 41 to 74%.
On BrowseComp-Plus, an existing evaluation set for research agents,[4] the dual-system agents cut active search time by 3.7 to 4.7 times and raise answer accuracy from 45% to as much as 54%. The phrase "up to" matters: this is the best result across the tested configurations, not a gain that every configuration achieved.
4. A concrete case: who makes which call in a safety literature search
Consider a search agent helping to collect candidate reports of an adverse event for a marketed drug from the published literature.
Inside a dual-system agent
System 2: I searched for the drug name together with the adverse event term. Please judge whether the top documents are relevant.
SearchJev: The first is a case report about the target drug, and the adverse event appears in the text. Relevant, high confidence. The second is a review of a different drug in the same class. Not relevant, also high confidence.
SearchJev: The third reports a combination regimen that includes the target drug, and the text does not say which drug caused the event. My confidence is low, so I am passing this one to System 2.
System 2: Keep the third as combination-therapy information and flag it for human review. The evidence is not yet sufficient, so I will change the query and search again.
What matters in this example is less the speed than the point at which a decision is handed off. If confidence tracks real accuracy, only low-confidence judgments need to go to a slower reasoner or to a person. If confidence cannot be trusted, the routing itself falls apart.
5. What is not new, and what the abstract does not settle
Handling most cases with a fast component and sending hard ones to a slower, more careful one is not a new idea. Cascades of light and heavy models, scoring models that rank candidates without generating text, and methods that calibrate predicted probabilities against observed accuracy have each been studied before. The contribution here, as the abstract presents it, is organizing the decisions inside a search agent by type into one benchmark, and combining calibration learned from uncertain labels with a hand-off for uncertain cases.
There are limits. First, the quality comparison is against same-size Qwen3.5 autoregressive models. The abstract does not say how SearchJev compares with larger models making the same decisions. Second, the calibration gains are measured on SearchDecision-Bench, which the authors built. Whether confidence stays reliable on questions or domains outside the training distribution is not shown.
Third, the speedup refers to active search time; how much end-to-end time to an answer shrinks is a separate question. Fourth, the abstract does not describe how the confidence threshold for handing off to System 2 is set, or how moving it shifts the balance between accuracy and speed. Fifth, it does not explain how the uncertain supervision used by SLCD was produced.
6. Why this topic is forming a cluster now
In this collection run, the paper shares a cluster with RoboQuest, for a cluster size of 2.[5] RoboQuest is a benchmark that asks whether agents that can move objects are able to search for, inspect, and test hidden information while doing a task. Its authors report that agents often stop exploring too early, making decisions before they have observed the evidence needed to complete the task.
One paper is about searching documents and the other about physical tasks, so the settings are very different. Both, however, touch the same question: how an agent decides that the evidence is sufficient. SearchJev aims to make that call quickly and with calibrated confidence; RoboQuest shows what goes wrong when the call is made too early.
On the paper-sharing site, SearchJev has 23 upvotes and 2 comments, and the repository with its training code has 15 stars.[2][3] These are signs of interest from researchers and implementers. They do not show that the claims are correct.
7. Where this connects to pharma and regulatory work
Pharma companies routinely do work that consists of picking relevant items out of large candidate sets and deciding whether the evidence is enough: safety literature surveillance, searches of regulatory documents, and responses to internal inquiries. In all of these, missing something relevant costs far more than reviewing something irrelevant.
Where this paper connects is less the speed than the routing: sending low-confidence judgments upward. Replace the upper tier with a human reviewer instead of a larger model, and the result is an operating model in which people see only the documents the machine is unsure about.
That routing is safe only if confidence is calibrated on the organization's own literature and questions. The authors' calibration results come from their own benchmark, and the work is a preprint. Before adopting anything like it, a team should check how confidence maps to actual accuracy on its own data, and set up a way to log and review cases where the model was confident and wrong.