Recursive AI development without independent verification. Input: R&D lead share surging from under 1% to 26% with roughly 30,000 agents. Recursive Closure: input and output stay in one organization. Unchecked Scaling: safety compute at roughly 6%, costs of intervention grow as decisions cascade. Structural Bias: self-assessment skewed by business stakes and closed methodology. Three Conditions: third-party log access, reproducible methods, and a safety compute floor. Output: verifiable R&D with outsider review. The frame wraps the flow with the principle of separating creator from reviewer, drawn from material review practice.
Image abstract — the whole article on one page (click to enlarge)

On September 17, 2026, Anthropic disclosed that its AI model Claude now leads 26 percent of its research and development, with roughly 30,000 AI agents running simultaneously. When AI leads a growing share of the work that builds its own successor, who verifies that the development is safe? Metrics published by the developer alone do not constitute verification; third-party access to logs and measurement methodology is required.

01Claude handled less than 1% of Anthropic's R&D just six months ago

Anthropic's blog post on September 17 contained three figures. The share of R&D where Claude "leads" — meaning the AI completes most of the work while a human supervises — reached 26 percent. The "collaborates" level, where AI participates in some capacity, exceeded 90 percent. And the number of agents running concurrently stood at roughly 30,000. Six months earlier, the lead figure was below 1 percent. The Washington Post described the jump as rising from zero in just months.

Figure 1 The rapid rise of Claude's R&D lead share
6 monthsShift toleadPresentFebruary 2026R&D lead shareunder 1%Collaborationover 90%Human instructs,AI executesAugust 2026R&D lead share26%30,000 agentsrunningContinuousdecisions6 monthsShift to leadPresentFebruary 2026R&D lead share under 1%Collaboration over 90%Human instructs, AI executesAugust 2026R&D lead share 26%30,000 agents runningContinuous decisions
AI lead share rose from near zero to 26% in six months, with roughly 30,000 agents running simultaneously.

The definition of "lead" matters. Anthropic defines it as the stage where a human supervises but the AI completes the majority of the work. This is not an AI writing boilerplate code under detailed instructions. It is an AI participating in design decisions. The shift from under 1 percent to 26 percent represents a qualitative change, not merely a quantitative one.

02When AI builds its successor, both input and output stay inside the same organization

In conventional software development, roles are distributed. Someone writes the specification, someone writes the code, and someone else verifies the result. Recursive AI development collapses this structure. The input to the development process is the model itself, and the output is the next model. Both remain within the same organization.

An external research institution that tries to reproduce the results has no access to the model's internal states or training data. Anthropic's decision to publish metrics is a meaningful step toward transparency, but no mechanism has been provided for outsiders to verify whether those metrics are measured correctly. Reuters' reporting on whether AI firms should be required to disclose dangerous incidents speaks directly to this structural gap.

03The longer recursive development proceeds without verification, the costlier it becomes to intervene

According to CNBC, Anthropic allocated approximately 6 percent of its total compute to safety work during a sample week in July. For research led by AI, that share rose to 12 percent — still just over a tenth of the total. With roughly 30,000 agents making decisions continuously, human oversight of every judgment has become physically impossible.

Once recursive development is underway, the cost of stopping it grows with time. When Model A's design decisions are embedded in Model B's training data, and Model B's outputs shape Model C, tracing the origin of a problem requires complete intermediate logs. Who stores those logs, in what format, and who may access them — none of this has been settled.

04Recursive AI development means the model participates in its own design decisions

A boundary is needed. AI involvement in development spans at least three stages.

StageHuman roleAI roleCurrent status
CollaborationIssues instructionsFollows instructionsOver 90%
LeadSupervisesCompletes most work independently26%
AutonomyNot involvedHandles entire process aloneNot yet reached
Figure 2 How recursive development closes off external verification
Recursive developmentAI builds its successorInput stays internalThe model itself is the inputMeasurement staysinternalThe company measures itselfVerification blockedOutsiders cannot reproduceRising correction costsDelay increases expenseRecursive developmentAI builds its successorInput stays internalThe model itself is the inputMeasurement stays internalThe company measures itselfVerification blockedOutsiders cannot reproduceRising correction costsDelay increases expense
In recursive AI development, both input (model) and measurement (metrics) remain within the same organization, making external verification structurally difficult.

The 26 percent Anthropic reports falls in the "lead" stage. Humans still supervise. But combined with the 90-plus percent at the collaboration level, the share of R&D work where AI plays no role at all has already fallen below 10 percent. The boundary of recursive AI development is not "AI writes code" but "AI participates in design decisions while humans only supervise."

05In material review, separating the creator from the reviewer is already institutionalized

1

The material review principle

The creator and the reviewer must be different people. Having the same individual perform both roles is not accepted under quality assurance standards.

2

The current state of recursive development

AI creates, the same company measures, and the same company reports "safe." The separation between creation and verification does not exist.

3

The shared principle

When the creator also serves as the verifier, reliability decreases regardless of intent.

In pharmaceutical material review, guidelines on promotional information activities require that companies establish review and oversight systems to ensure the appropriateness of materials. The person who drafts a promotional material does not review it; a different individual does. This separation is necessary even when the creator acts in good faith. Regardless of intent, when the same party handles both creation and verification, an optimistic bias arises structurally. The same principle applies to AI development.

06Self-assessment bias is structural, not motivational

1

Aligned incentives

When assessment results directly affect business continuity, an optimistic bias emerges.

2

Information asymmetry

When outsiders cannot examine internal logs and methodology, the validity of published figures cannot be verified.

Anthropic's publication of metrics is a progressive move within the industry. Few other AI companies have disclosed comparable figures. But the issue is not whether good intentions exist. When evaluation outcomes directly influence fundraising, regulatory positioning, and market trust, an optimistic bias arises structurally. The bias at issue here, however, is not an individual cognitive tendency but one generated by the organization's incentive structure.

TechCrunch's question — whether the AI safety debate is really about "safety" or about "control" — approaches this structure from another angle. When a company defines safety, measures it, and declares it achieved, the discussion is simultaneously about safety and about who gets to set the pace of AI development.

07Without third-party access to unedited logs, the safety of recursive development cannot be confirmed

Making recursive development externally verifiable requires three conditions.

Figure 3 Three conditions for verifiable recursive development
MethodologicaltransparencyResourceassuranceUnedited log accessThird-party reviewPre-publishedmethodsSafety computefloorExternally monitoredratioMethodologicaltransparencyResource assuranceUnedited log accessThird-party reviewPre-published methodsSafety compute floorExternally monitored ratio
Only when log access, measurement reproducibility, and a safety compute floor are all in place can recursive development be verified from outside.

First, third-party access to unedited development logs. Which agent made which decision, and how that decision was incorporated into the next model, must be reviewable by an institution other than the developer. If logs are provided in an editable state, the value of verification diminishes sharply.

Second, advance publication and reproducibility of measurement methods. The definitions behind "26 percent" and "over 90 percent" must be published in advance and be reproducible by independent researchers.

Third, a minimum threshold for safety compute ratios. Whether the 6 percent figure from a sample week is a constant allocation remains unknown. Even if AI-led research raises it to 12 percent, who monitors the ratio and what happens when it drops below the threshold has not been decided.

In material review practice, this structure translates to a straightforward operational rule: do not let the same AI that drafted a material verify it. Asking AI whether its own output is acceptable is not verification. In an era where Bloomberg Law reports that runaway AI agents are now real, the independence of the verifier becomes a prerequisite for institutional design.

  1. Anthropic's Claude grew from under 1% to 26% of R&D leadership in six months, with roughly 30,000 agents running simultaneously, but these figures are self-reported using the company's own measurements
  2. In recursive development, input, output, and measurement all stay within the same organization, so safety claims cannot be confirmed without independent external verification
  3. The material review principle of separating creator from reviewer applies equally to AI development — the independence of the verifier is the condition for trust
Closing

In an era where AI builds its own successor, the developer's self-report cannot be the sole evidence. Anthropic's decision to publish metrics is a step forward, but without the means for outsiders to verify whether those metrics are measured correctly, transparency remains superficial. Third-party access to logs and measurement methodology is the first condition for confirming the safety of recursive development.

  1. The Washington Post, "Anthropic says its chatbot Claude is taking over the work of building its own successor," September 17, 2026. Link
  2. Anthropic, "Measurements for understanding the pace of AI development inside frontier labs," September 17, 2026. Link
  3. CNBC, "Anthropic shares 3 metrics to help AI companies monitor pace of development," September 17, 2026. Link
  4. Reuters, "Should AI firms be required to disclose dangerous incidents?" September 16, 2026. Link
  5. Technology.org, "Claude Now Leads 26% of Anthropic's AI Research," September 18, 2026. Link
  6. Ministry of Health, Labour and Welfare (Japan), "Guidelines on Promotional Information Activities for Prescription Drugs," September 25, 2018. Link
  7. Bloomberg Law, "Runaway AI Agents Are Now Real. Companies Need to Rethink Governance," September 17, 2026. Link
  8. TechCrunch, "Is the AI safety debate really about 'safety' or 'control'?" September 17, 2026. Link