A research request first goes to several planning agents, which explore external sources and build a plan called a ResearchSpec. Research agents then each take a section and investigate and draft it in parallel. Finally the assembled report gets a global review, and only the affected parts are revised. The same workflow also builds research tasks and trajectories for LongCat's mid-training and post-training. Combining planning perspectives helped, but plan refinement and readability effects were mixed, and the evaluation is by the developers in a preprint.
Image abstract — the whole article on one page (click to enlarge)

1. Where long AI-written research reports tend to break

Searching the web and the literature, then assembling a long report with evidence attached: this "deep research" use of language models is becoming routine. But the wider the scope and the longer the report, the harder it is for one model working in one context to handle planning, investigation and writing all at once.

Two failure modes are common. The original plan blurs as the investigation proceeds, so sections end up uneven in depth and angle. And when only one part needs fixing, the whole report gets rewritten, changing sections that were already fine.

This technical report from Meituan's LongCat team describes a deep research system that pairs an enhanced LongCat model with a multi-agent workflow. The primary source is a preprint on arXiv and has not been peer reviewed. As a technical report, it is also an evaluation by the system's own developers, which is worth keeping in mind while reading.[1]

2. The core idea: separate global planning from section-level investigation

Reduced to one design choice, the system splits planning for the whole report from detailed investigation of each section, and carries out revisions at the section level.

The workflow runs in three stages. First, several planning agents explore external sources and refine an actionable research plan, which the authors call a ResearchSpec. Next, research agents each take an assigned section, investigate it and draft it in parallel, gathering extra evidence in separate contexts as their analyses develop. Finally, once the sections are assembled, a global review identifies problems and guides targeted local revisions. According to the authors, this reduces reliance on repeated full-report rewriting.

The authors also use the same workflow to build research tasks and trajectories for the mid-training and post-training of LongCat's general-purpose models. Running the system doubles as a way of producing training data for the next model.

3. What the abstract reports

On public benchmarks, the system is reported to score 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II and 79.83 on ResearchRubrics. On an in-house benchmark it scores 76.04, ranking second among four compared systems.[4]

Analyses on a development set show benefits from combining planning perspectives, while further refining the plan had mixed effects. Adding a final editing step improved average automatic readability preference across two benchmarks, but the trend differed between them.

Those hedged phrases deserve attention. Technical reports tend to lead with good numbers, yet this abstract also mentions what did not clearly help. A careful reading separates which design choices worked, and under which conditions.

4. A worked example: a report on public information in one disease area

Suppose a team needs a single report on one disease area covering treatment options, published results from major clinical trials, and recent actions by regulators in several countries. The scope is broad, and each section draws on entirely different sources.

One context writes everything

The model investigates and writes in a single pass, then reviews the whole. When an error turns up in the regulatory section, fixing it means rewriting the entire report, and the wording and citations in the clinical trial section, which had no problems, change as well. Tracking what changed becomes hard for the people checking it.

Planning and section research are separated

A research plan is written first, stating what each section will cover and from which angle. Separate agents investigate and draft each section. When the global review finds an error in the regulatory section, only that section is revised. Because the plan exists as a document, a reviewer can later check what was supposed to be investigated.

The benefit of the second approach is less about whether the output is correct and more about being able to trace what was changed and why. Inconsistencies across sections, such as two sections describing the same trial differently, still depend on the global review catching them. The abstract does not say how well it does that.

5. What is not new, and what the abstract leaves open

Planning first, researching from several perspectives, and then assembling a long text is not a new idea. Earlier work on generating Wikipedia-like articles from scratch already used perspective-guided questioning to research a topic and build an outline before writing.[5] Multi-agent role splitting is also common. What this report adds is the combination of parallel section-level research with local revision, tied to training data generation.

The limits are considerable. First, public benchmark scores are reported, but the abstract names neither the systems compared nor the size of any gap. Second, on the in-house benchmark the system ranks second, not first, and that benchmark cannot be checked from outside. Third, readability is measured by automatic preference, which is not the same as whether a human reader finds the text clear. Fourth, for reports described as evidence-grounded, the abstract does not say how the authors verified that cited sources actually support the claims. Fifth, compute cost and running time are not given.

The public repository presents the work as a harness: users connect their own language model, web search and page-fetching backends, and the documentation states that outputs need human review.[2] The reported scores come from pairing the workflow with the enhanced LongCat model, so plugging in a different model will not necessarily reproduce them.

6. Why this report surfaced now, and why it is not part of a cluster

The report received 66 upvotes on Hugging Face Daily Papers, and its public repository has 18 stars.[3] Upvotes are reader votes and stars are votes from people interested in the code; both measure attention. Neither says whether the report is correct or whether the system is better than others.

In this site's selection, the only signal that fired was reader votes. The star count did not reach the level treated as implementation follow-through, and no press coverage was found. No cluster of related papers formed either; the cluster size is 1. Within this week's collection, no group of papers converged on the same question as this report.

The report is best read as one worked case of how to structure a deep research workflow. It is not evidence that the topic as a whole is moving all at once.

7. Where it connects to pharmaceutical and regulatory work

Pharmaceutical work involves many tasks that require broad investigation followed by an evidence-backed document: surveying a disease area, tracking regulatory developments, compiling safety literature. If deep research systems are used for these, parts of this report's design are worth borrowing.

One is writing the research plan down as a document before investigating. If what will be examined, from which angle and to what depth is recorded, reviewers can check whether the plan is sound before looking at conclusions. That fits well with defining investigation scope in standard operating procedures. Another is keeping revisions at the section level, which prevents approved sections from being rewritten by accident. Teams that care about version control and change history will find that easier to manage than whole-document rewrites.

Whatever the design, a person still has to confirm that cited evidence supports each claim. This is a preprint, and its scope does not establish whether the same performance holds for medical or regulatory documents. Reading the in-house benchmark result as a proxy for performance on a company's own tasks would also be a mistake.