
1. How do you train an agent that works with real files?
What people expect from language model agents is shifting from answering questions to doing actual work: reading spreadsheets, consulting PDF reports, switching between tools and producing a finished deliverable. The authors call these "working agents."
Training such agents requires large numbers of tasks built on many realistic files, with results that can be verified. The authors note that few pipelines produce this kind of data. Existing approaches have one of two weaknesses. Some have a model generate the files, which then lack realism and diversity. Others build tasks on real files but without task-specific verifiers, so the quality of an agent's output goes unchecked.[1]
This paper proposes a data synthesis framework that uses real files and ties both the task and its grading to those files. The primary source is a preprint on arXiv and has not been peer reviewed.
2. The core idea: derive the task and its rubric from the same evidence graph
Reduced to one idea, GraphForge builds an evidence graph over the relationships among real files gathered into a workspace, and derives both the task statement and its grading rubric from that graph.
The process goes like this. It starts from occupation-grounded seeds, which keep diversity under control. For each seed, it assembles a workspace of real files and builds an evidence graph over how those files relate. Because the task statement and the rubric both come from this graph, the task's requirements are backed by files in the workspace, and each rubric criterion is anchored to the files needed to verify it.
Before collecting training trajectories, an initial rollout tests whether the task can actually be executed. If something is wrong, a revision agent repairs the task and rubric against the original files. Building checking and repair into the data creation stage is the distinctive part.
3. What the abstract reports
The authors fine-tuned Qwen3.6-27B on 2,169 GraphForge trajectories. Running under OpenHands, the model reached 1445.7 on GDPVal, an increase of 65.7.[3][5] Running under Claude Code, it scored 63.7 on Workspace-Bench-Lite (up 7.7) and 24.0 on SpreadsheetBench II (up 13.7).[4]
The authors also tried rejection fine-tuning: the fine-tuned model generated its own rollouts, candidates were selected using the evidence-anchored rubrics, and the model was trained again on the selected ones. This produced further gains on all three benchmarks, which the authors take as a sign that the rubrics provide a useful selection signal.
The abstract says the data and models are available. On the arXiv page itself, however, no link to where they are hosted could be found.
4. A worked example: finding what to change in materials after a label revision
Consider turning one pharmaceutical task into a training task. The package insert for a product has been revised, and someone needs to identify which parts of existing information materials must be updated and produce a list of the required changes.
Model-generated files
The model writes the old and new package inserts and the materials. They look tidy, but they rarely carry the friction of real documents: idiosyncratic phrasing, broken tables, versions of different ages mixed together. Grading tends to stop at "was a list produced," and it is hard to check whether each listed item really corresponds to the revision.
Real files and an evidence graph
Real documents are placed in the workspace, and relationships such as "this statement in this material corresponds to this revised section of the package insert" are built into a graph. The task statement is derived from those relationships, and each grading criterion is linked to which file, and where in it, can confirm the answer. Each change the agent proposes can be checked against its source file.
The second approach keeps grading criteria in a form that can be traced back to source documents, which is close to how people cross-check work. But this is about how training data is built. Whether an agent trained this way identifies the right changes in real work has to be tested separately.
5. What is not new, and what the abstract leaves open
Evaluating models on occupation-grounded, real-world tasks was already the idea behind GDPval.[3] Benchmarks for practical spreadsheet manipulation exist too.[4] Selecting candidates with a rubric and retraining on them is also a known technique. What this paper adds, on its own account, is deriving the task and rubric from one graph over real files, with every criterion anchored to specific files.
There are several limits. First, only one model was fine-tuned, so the abstract does not show whether other sizes or model families benefit in the same way. Second, the abstract does not explain the GDPVal scale or what an increase of that size means in practice. Third, gains are reported on Workspace-Bench-Lite and SpreadsheetBench II, but the abstract does not say whether they were compared against training on the same amount of other data. Fourth, the same framework both creates the rubrics and uses them to select training candidates, so any bias in the rubrics could feed straight into training. Fifth, the abstract does not say where the real files came from or how rights and personal data were handled.
6. Why this paper surfaced now, and why it is not part of a cluster
The paper drew 144 upvotes on Hugging Face Daily Papers,[2] the most among today's three selections. Upvotes are reader votes on what seems worth reading; they do not show that the claims are correct or that the method is better than alternatives.
In this site's selection, the only signal that fired was reader votes. At collection time no public repository was registered, the star count was 0, and no press coverage was found. No cluster of papers on the same question formed; the cluster size is 1.
How to build training data for agents has come up repeatedly on this site in recent days, in papers on agent task authoring and action verification. None of those fell into the same cluster as this paper, though. The accurate reading is not that the topic is moving all at once, but that papers with related questions are appearing separately.
7. Where it connects to pharmaceutical and regulatory work
Document work in pharmaceutical companies looks very much like what the paper calls working-agent tasks. People read across package inserts, study reports, promotional materials and internal procedures, then produce a deliverable in a fixed format. When training or evaluating agents for that work, a few ideas from this paper carry over.
One is tying each grading criterion to the source document used to check it. If agent output is judged by "correct against which document, and where" rather than "looks plausible," the basis for the evaluation can be explained in an audit. Another is running each task once after creating it and fixing tasks that cannot be executed before using them, which applies directly to building an internal evaluation set.
Using real internal documents as training data raises separate issues of confidentiality and personal information. The abstract does not reveal what kind of files the paper used as public data, so anyone considering internal use needs to settle that first. This is a preprint, and none of the effects described here have been tested on pharmaceutical document work.