Flow for a result nobody asked for. First, scope: Claude showed at least 67.2% of zeros lie on the line, while the hypothesis stays open. Second, sort the checks: the 13 validators: Claude itself, while Lean and a separate proof are outside. Third, ask for the trail, such as the 650 failed ideas. Last, wait for journal peer review and use the result conditionally until then.
Image abstract — the whole article on one page (click to enlarge)

On 1 October 2026, Live Science reported that a research version of Anthropic's AI model Claude had failed to prove the Riemann hypothesis, and had produced a different mathematical result along the way. When an AI hands back a result nobody asked for, what does it have to pass before we can call it correct? My answer: checks that are independent of the AI. Claude's result has passed a machine check in the proof assistant Lean and a separate proof by the mathematician Youness Lamzouri. It has not yet passed journal peer review.

01The AI failed at the Riemann hypothesis and proved a different theorem on the way

Anthropic asked the research version of Claude to prove the Riemann hypothesis. Claude failed. What Claude returned instead was a result nobody had requested.

The Riemann hypothesis concerns the points where a function called the zeta function takes the value zero. Mathematicians call these points zeros. The hypothesis says that all the non-trivial zeros lie on a single line, the line where the real part equals 1/2. For decades mathematicians have been proving larger and larger lower bounds on the share of zeros that sit on that line. According to Anthropic, the best proven share had been 41.6%. Claude proved that at least 67.2% of the zeros lie on the line.

Figure 1 From a failed request to an unrequested result
Asked to provethe hypothesisAnthropic's request650 ideas failFirst sessionAbout 60subagentsSecond sessionBound rises to67%An unrequestedresultAsked to prove the hypothesisAnthropic's request650 ideas failFirst sessionAbout 60 subagentsSecond sessionBound rises to 67%An unrequested result
The request was for a proof of the hypothesis; what came back was a different result, produced in the middle of the failure.

Anthropic published the result on its own website on 10 August 2026. Since then Claude's result has gone through several checks. The proof assistant Lean checked the proof by machine. Two mathematicians at Anthropic read the proof, and the paper was sent to two outside experts. On 2 September, Lamzouri proved the same share by a different route. Journal peer review has not yet happened.

My conclusion is simple. Whether to accept an AI's result does not depend on how impressive the result is. It depends on whether something or someone outside the AI has checked it. The 13 validator agents that ran inside Claude do not count as an outside check.

02Claude raised a lower bound; the hypothesis itself remains open

Claude returned a result about a share of zeros, not the proof it was asked for. Before going further, it is worth separating what that result shows from what it does not.

The Clay Mathematics Institute lists the Riemann hypothesis as one of its prize problems. According to the Institute, the first ten trillion zeros have been checked by computation and lie on the line. There are infinitely many zeros, so computation alone cannot prove the hypothesis.

AspectRiemann hypothesisThis result
ClaimEvery non-trivial zero has real part 1/2More than 67% of zeros lie on the line
StatusUnsolved. A Clay prize problemPreprint. Not yet peer reviewed
ChecksFirst ten trillion zeros confirmed by computationLean, two mathematicians at Anthropic, a separate proof

The hypothesis is a claim about every zero. Claude's result says only that more than 67% of the zeros lie on the line. It tells us nothing about where the remaining third or so of the zeros lie. Live Science also reported that the result is not progress on the hypothesis itself.

Anthropic wrote that it does not expect Claude's method to lead to a proof of the hypothesis. The company that announced the result has itself limited how far the result reaches.

The sources give two slightly different figures. Anthropic's announcement gives 67.2%; Lamzouri's abstract and the Live Science article give 67.25%. In this piece I say which source each figure comes from.

03The checks came from Claude's validators, Lean, two Anthropic mathematicians and a separate prover

Claude's result is not a proof of the hypothesis, but it is still a new theorem. Here, in order, is who checked that new theorem and where.

The first check took place inside Claude. During the work, Claude handed the role of validator to 13 smaller agents. The validator agents looked for counterexamples and redid the proof from scratch.

Next, Claude's proof was checked in Lean. Lean is software in which a computer checks a mathematical proof line by line. To check a proof in Lean, the proof first has to be rewritten in a form Lean can read.

Human mathematicians read Claude's proof too. Alpöge and Furman, mathematicians at Anthropic, checked it. Lamzouri's abstract also states that Claude's proof was checked by these two. According to Channel Insider, Anthropic also shared the paper with two outside experts, Conrey and Goldston.

Finally, on 2 September, Lamzouri posted a different proof on arXiv, the preprint server. Claude's proof had been built on a framework that uses matrices. Lamzouri did without the matrices and reached the same share through a different kind of inequality, set in a Hilbert space.

Who checkedIndependent of the AI?What they examined
13 validator agentsNo; they are ClaudeSearches for counterexamples, re-proofs from scratch
LeanYes (a small trusted kernel)The formalised proof
Alpöge, Furman, Conrey, GoldstonTwo inside Anthropic, two outsideThe paper
LamzouriYes (a separate mathematician)A different proof without matrices

Anthropic posted Claude's paper on its own website, not on arXiv. Channel Insider noted that the result had not gone through the normal academic peer-review process.

04Validators built from the same model are not independent of it

Five parties checked Claude's result or received the paper: the validator agents, Lean, two mathematicians inside Anthropic, two experts outside, and Lamzouri. I want to sort out which of them are independent of Claude.

The 13 validators that first examined Claude's result were all agents run by Claude itself. Whether they hunted for counterexamples or redid the proof, the judgement was made inside the same model. So I do not count the 13 validators as a check independent of the AI. I am not saying the validators missed anything. I am saying only that Claude checking its own answer is not a check from outside.

Lean sits outside Claude. The Lean website explains that a minimal trusted kernel guarantees the correctness of a proof. Lean compares the rewritten proof, line by line, against its rules, regardless of how Claude arrived at it.

Lamzouri's proof is also independent of Claude. Lamzouri reached the same share as Claude without using matrices. If there were an error somewhere in Claude's proof, Lamzouri's proof, taking a different route, would not inherit it.

Figure 2 Checks inside the AI and checks outside it
Inside theAIOutside theAI13 validatorsClaude itselfCounterexamplesearchInside checkOwn re-proofInside checkLean checkMinimal trusted kernelFour mathematiciansRead the paperA different proofLamzouriInside the AIOutside the AI13 validatorsClaude itselfCounterexamplesearchInside checkOwn re-proofInside checkLean checkMinimal trusted kernelFourmathematiciansRead the paperA different proofLamzouri
The upper row is checking within the same model; the lower row is checking independent of the AI. Acceptance rests on the lower row.

Alpöge and Furman are human readers who read the proof separately from Claude. But they also work for Anthropic, the company that announced the result. I have not been able to find out what the outside experts, Conrey and Goldston, concluded.

05Of about 60 subagents, two produced the key ideas

The progress of Claude's work was kept as a record that could be counted. I think that record gives outside readers a way to follow Claude's work. Here is what Anthropic published.

According to Anthropic's report, Claude used 31 million output tokens over two sessions. All 650 ideas Claude tried in the first session failed.

In the second session, Claude ran about 60 smaller agents. Anthropic published, in numbers, what those agents did.

1

Produced the key ideas

Two. These two agents built the central mathematical ideas.

2

Passed on ideas

Thirteen. These agents sent ideas to the central two.

3

Found nothing new

Thirty. These agents tried to develop new ideas and could not.

4

Validated

Thirteen checked the work; another two helped draft the paper.

In the second session, 2,400 shell commands were run. Anthropic did not show only the two agents that produced the key ideas. It also counted and published the 30 agents that produced nothing.

An outside mathematician reading Anthropic's record can follow where Claude tried what, and where it got stuck. I think the fact that Anthropic published even the number of failures helps the people checking from outside.

06Independent checks, a recorded trail, and the time peer review takes

The 650 failed ideas and the 2,400 commands were on record, so outside mathematicians can follow the work. From that, I draw three grounds for accepting a result an AI hands back.

An agent's "verified" is not grounds for acceptance

Claude's 13 validators were Claude. The checks from outside Claude came from Lean and Lamzouri. When an AI agent reports that its work has been verified, the person receiving it cannot accept the result on that report alone. The person receiving it will need to ask whether someone or something outside the AI has checked it.

Ask for the trail along with the result

Anthropic published the 650 failed ideas and a breakdown of what the 60 agents did. Outside mathematicians can use that record to follow the work. Anyone who hands work to an AI should also collect a record of what it tried and where it failed, not only the finished output. Without such a record, an outsider faced with a bare result has to check everything again from the start.

The day a machine check passes and the day mathematicians accept a result are different days

The Lean check had been completed by the time Anthropic announced the result on 10 August. Lamzouri's proof appeared a little over three weeks later. Journal peer review has not finished. The Clay Institute's rules for its prizes require at least two years to pass after publication, and general acceptance in the mathematical community. Those rules are written for solving the hypothesis itself, but they set community acceptance as a condition separate from any machine check. The day Lean accepts a proof and the day mathematicians accept a result are different days.

Figure 3 Before accepting a result nobody asked for
not yetReceivethe…Ask for thetrail tooCheckedindependently?MachinecheckWhereformalisableOutsidersread itReproduceanother…Use withconditionsUntil peerreviewnot yetReceive the resultAsk for the trail tooCheckedindependently?Machine checkWhere formalisableOutsiders read itReproduce another wayUse with conditionsUntil peer review
Stop an agent's own 'verified' report at the decision point. Even after machine and human checks, treat the result as conditional until peer review.

Claude's result has passed checks from outside the AI, and the record of the work has been published. Only journal peer review remains.

07Until journal review is done, Claude's 67% is still awaiting acceptance by mathematicians

What remains for Claude's result is journal peer review. Here is what might happen while that review is under way.

Both Claude's paper and Lamzouri's paper are preprints, not yet reviewed. How journal referees will judge them is not yet known. I am not going to predict the outcome.

1

Journal review

Neither Claude's paper nor Lamzouri's has a peer-review verdict yet.

2

The path to the hypothesis

Anthropic wrote that it does not expect Claude's method to solve the hypothesis.

Claude failed to prove the Riemann hypothesis, and raised the bound on the share of zeros while trying. The figure of 67% did not come from a shortcut toward the hypothesis.

Another question is still open: how journal review will keep up with the speed at which AI produces mathematical results. This time, a little over three weeks separated Anthropic's announcement from Lamzouri's separate proof. The Clay rules work on a timescale of years. How long peer review will take cannot be told from the sources used here. No published material yet shows whether that gap between AI and peer review will narrow or widen.

Key Points ── 3 to take away
  1. The 13 validators that first examined Claude's result were Claude itself; the outside checks came from the proof assistant Lean and the mathematician Lamzouri. An agent saying "verified" is not grounds for acceptance.
  2. Anthropic published the 650 failed ideas and what the 60 agents did, so outside mathematicians can follow the work. Ask for the trail along with the result.
  3. Claude's result has not been peer reviewed. The Clay Institute's rules require two years after publication and general acceptance. Treat fast checks and slow checks as different things.
Closing

When an AI returns a result nobody asked for, we can call it correct only after it has passed checks from outside the AI. Validation carried out inside the same AI is not such a check.

Claude's 67% has passed a machine check in Lean and a separate proof by Lamzouri. What remains is the slower check of journal peer review. A Lean check and a separate proof finishing early do not mean peer review has been done.

Sources & references
  1. Live Science. 'This is what happened with Claude and me': AI's failed attempt to crack the Riemann hypothesis led mathematician to a breakthrough. 2026-10-01.(Report of the event; Lamzouri's role and the 67.25% figure)
  2. Anthropic. Learning more about Claude's mathematical capabilities. 2026-08-10.(The improvement from 41.6% to 67.2%, the 650 failed ideas, the agent breakdown, checks by Lean and mathematicians)
  3. Youness Lamzouri. A new proof that more than 2/3 of the zeros of the Riemann zeta function are simple and on the critical line. arXiv, 2026-09-02.(A different proof without matrices)
  4. Channel Insider. Unreleased Claude Model Makes Breakthrough on 167-Year-Old Math Problem. 2026-08-17.(Not yet peer reviewed; the paper shared with outside experts)
  5. Clay Mathematics Institute. Riemann Hypothesis. Accessed 2026-10-03.(Statement of the hypothesis and the first ten trillion zeros)
  6. Lean FRO. Lean: Programming Language and Theorem Prover. Accessed 2026-10-03.(Proof checking by a minimal trusted kernel)
  7. Clay Mathematics Institute. Rules for the Millennium Prize Problems. Accessed 2026-10-03.(Two years after publication and general acceptance as prize conditions)