
On 1 October 2026, Live Science reported that a research version of Anthropic's AI model Claude had failed to prove the Riemann hypothesis, and had produced a different mathematical result along the way. When an AI hands back a result nobody asked for, what does it have to pass before we can call it correct? My answer: checks that are independent of the AI. Claude's result has passed a machine check in the proof assistant Lean and a separate proof by the mathematician Youness Lamzouri. It has not yet passed journal peer review.
01The AI failed at the Riemann hypothesis and proved a different theorem on the way
Anthropic asked the research version of Claude to prove the Riemann hypothesis. Claude failed. What Claude returned instead was a result nobody had requested.
The Riemann hypothesis concerns the points where a function called the zeta function takes the value zero. Mathematicians call these points zeros. The hypothesis says that all the non-trivial zeros lie on a single line, the line where the real part equals 1/2. For decades mathematicians have been proving larger and larger lower bounds on the share of zeros that sit on that line. According to Anthropic, the best proven share had been 41.6%. Claude proved that at least 67.2% of the zeros lie on the line.
Anthropic published the result on its own website on 10 August 2026. Since then Claude's result has gone through several checks. The proof assistant Lean checked the proof by machine. Two mathematicians at Anthropic read the proof, and the paper was sent to two outside experts. On 2 September, Lamzouri proved the same share by a different route. Journal peer review has not yet happened.
My conclusion is simple. Whether to accept an AI's result does not depend on how impressive the result is. It depends on whether something or someone outside the AI has checked it. The 13 validator agents that ran inside Claude do not count as an outside check.
02Claude raised a lower bound; the hypothesis itself remains open
Claude returned a result about a share of zeros, not the proof it was asked for. Before going further, it is worth separating what that result shows from what it does not.
The Clay Mathematics Institute lists the Riemann hypothesis as one of its prize problems. According to the Institute, the first ten trillion zeros have been checked by computation and lie on the line. There are infinitely many zeros, so computation alone cannot prove the hypothesis.
| Aspect | Riemann hypothesis | This result |
|---|---|---|
| Claim | Every non-trivial zero has real part 1/2 | More than 67% of zeros lie on the line |
| Status | Unsolved. A Clay prize problem | Preprint. Not yet peer reviewed |
| Checks | First ten trillion zeros confirmed by computation | Lean, two mathematicians at Anthropic, a separate proof |
The hypothesis is a claim about every zero. Claude's result says only that more than 67% of the zeros lie on the line. It tells us nothing about where the remaining third or so of the zeros lie. Live Science also reported that the result is not progress on the hypothesis itself.
Anthropic wrote that it does not expect Claude's method to lead to a proof of the hypothesis. The company that announced the result has itself limited how far the result reaches.
The sources give two slightly different figures. Anthropic's announcement gives 67.2%; Lamzouri's abstract and the Live Science article give 67.25%. In this piece I say which source each figure comes from.
03The checks came from Claude's validators, Lean, two Anthropic mathematicians and a separate prover
Claude's result is not a proof of the hypothesis, but it is still a new theorem. Here, in order, is who checked that new theorem and where.
The first check took place inside Claude. During the work, Claude handed the role of validator to 13 smaller agents. The validator agents looked for counterexamples and redid the proof from scratch.
Next, Claude's proof was checked in Lean. Lean is software in which a computer checks a mathematical proof line by line. To check a proof in Lean, the proof first has to be rewritten in a form Lean can read.
Human mathematicians read Claude's proof too. Alpöge and Furman, mathematicians at Anthropic, checked it. Lamzouri's abstract also states that Claude's proof was checked by these two. According to Channel Insider, Anthropic also shared the paper with two outside experts, Conrey and Goldston.
Finally, on 2 September, Lamzouri posted a different proof on arXiv, the preprint server. Claude's proof had been built on a framework that uses matrices. Lamzouri did without the matrices and reached the same share through a different kind of inequality, set in a Hilbert space.
| Who checked | Independent of the AI? | What they examined |
|---|---|---|
| 13 validator agents | No; they are Claude | Searches for counterexamples, re-proofs from scratch |
| Lean | Yes (a small trusted kernel) | The formalised proof |
| Alpöge, Furman, Conrey, Goldston | Two inside Anthropic, two outside | The paper |
| Lamzouri | Yes (a separate mathematician) | A different proof without matrices |
Anthropic posted Claude's paper on its own website, not on arXiv. Channel Insider noted that the result had not gone through the normal academic peer-review process.
04Validators built from the same model are not independent of it
Five parties checked Claude's result or received the paper: the validator agents, Lean, two mathematicians inside Anthropic, two experts outside, and Lamzouri. I want to sort out which of them are independent of Claude.
The 13 validators that first examined Claude's result were all agents run by Claude itself. Whether they hunted for counterexamples or redid the proof, the judgement was made inside the same model. So I do not count the 13 validators as a check independent of the AI. I am not saying the validators missed anything. I am saying only that Claude checking its own answer is not a check from outside.
Lean sits outside Claude. The Lean website explains that a minimal trusted kernel guarantees the correctness of a proof. Lean compares the rewritten proof, line by line, against its rules, regardless of how Claude arrived at it.
Lamzouri's proof is also independent of Claude. Lamzouri reached the same share as Claude without using matrices. If there were an error somewhere in Claude's proof, Lamzouri's proof, taking a different route, would not inherit it.
Alpöge and Furman are human readers who read the proof separately from Claude. But they also work for Anthropic, the company that announced the result. I have not been able to find out what the outside experts, Conrey and Goldston, concluded.
05Of about 60 subagents, two produced the key ideas
The progress of Claude's work was kept as a record that could be counted. I think that record gives outside readers a way to follow Claude's work. Here is what Anthropic published.
According to Anthropic's report, Claude used 31 million output tokens over two sessions. All 650 ideas Claude tried in the first session failed.
In the second session, Claude ran about 60 smaller agents. Anthropic published, in numbers, what those agents did.
Produced the key ideas
Two. These two agents built the central mathematical ideas.
Passed on ideas
Thirteen. These agents sent ideas to the central two.
Found nothing new
Thirty. These agents tried to develop new ideas and could not.
Validated
Thirteen checked the work; another two helped draft the paper.
In the second session, 2,400 shell commands were run. Anthropic did not show only the two agents that produced the key ideas. It also counted and published the 30 agents that produced nothing.
An outside mathematician reading Anthropic's record can follow where Claude tried what, and where it got stuck. I think the fact that Anthropic published even the number of failures helps the people checking from outside.
06Independent checks, a recorded trail, and the time peer review takes
The 650 failed ideas and the 2,400 commands were on record, so outside mathematicians can follow the work. From that, I draw three grounds for accepting a result an AI hands back.
An agent's "verified" is not grounds for acceptance
Claude's 13 validators were Claude. The checks from outside Claude came from Lean and Lamzouri. When an AI agent reports that its work has been verified, the person receiving it cannot accept the result on that report alone. The person receiving it will need to ask whether someone or something outside the AI has checked it.
Ask for the trail along with the result
Anthropic published the 650 failed ideas and a breakdown of what the 60 agents did. Outside mathematicians can use that record to follow the work. Anyone who hands work to an AI should also collect a record of what it tried and where it failed, not only the finished output. Without such a record, an outsider faced with a bare result has to check everything again from the start.
The day a machine check passes and the day mathematicians accept a result are different days
The Lean check had been completed by the time Anthropic announced the result on 10 August. Lamzouri's proof appeared a little over three weeks later. Journal peer review has not finished. The Clay Institute's rules for its prizes require at least two years to pass after publication, and general acceptance in the mathematical community. Those rules are written for solving the hypothesis itself, but they set community acceptance as a condition separate from any machine check. The day Lean accepts a proof and the day mathematicians accept a result are different days.
Claude's result has passed checks from outside the AI, and the record of the work has been published. Only journal peer review remains.
07Until journal review is done, Claude's 67% is still awaiting acceptance by mathematicians
What remains for Claude's result is journal peer review. Here is what might happen while that review is under way.
Both Claude's paper and Lamzouri's paper are preprints, not yet reviewed. How journal referees will judge them is not yet known. I am not going to predict the outcome.
Journal review
Neither Claude's paper nor Lamzouri's has a peer-review verdict yet.
The path to the hypothesis
Anthropic wrote that it does not expect Claude's method to solve the hypothesis.
Claude failed to prove the Riemann hypothesis, and raised the bound on the share of zeros while trying. The figure of 67% did not come from a shortcut toward the hypothesis.
Another question is still open: how journal review will keep up with the speed at which AI produces mathematical results. This time, a little over three weeks separated Anthropic's announcement from Lamzouri's separate proof. The Clay rules work on a timescale of years. How long peer review will take cannot be told from the sources used here. No published material yet shows whether that gap between AI and peer review will narrow or widen.
- The 13 validators that first examined Claude's result were Claude itself; the outside checks came from the proof assistant Lean and the mathematician Lamzouri. An agent saying "verified" is not grounds for acceptance.
- Anthropic published the 650 failed ideas and what the 60 agents did, so outside mathematicians can follow the work. Ask for the trail along with the result.
- Claude's result has not been peer reviewed. The Clay Institute's rules require two years after publication and general acceptance. Treat fast checks and slow checks as different things.
When an AI returns a result nobody asked for, we can call it correct only after it has passed checks from outside the AI. Validation carried out inside the same AI is not such a check.
Claude's 67% has passed a machine check in Lean and a separate proof by Lamzouri. What remains is the slower check of journal peer review. A Lean check and a separate proof finishing early do not mean peer review has been done.
- Live Science. 'This is what happened with Claude and me': AI's failed attempt to crack the Riemann hypothesis led mathematician to a breakthrough. 2026-10-01.(Report of the event; Lamzouri's role and the 67.25% figure)
- Anthropic. Learning more about Claude's mathematical capabilities. 2026-08-10.(The improvement from 41.6% to 67.2%, the 650 failed ideas, the agent breakdown, checks by Lean and mathematicians)
- Youness Lamzouri. A new proof that more than 2/3 of the zeros of the Riemann zeta function are simple and on the critical line. arXiv, 2026-09-02.(A different proof without matrices)
- Channel Insider. Unreleased Claude Model Makes Breakthrough on 167-Year-Old Math Problem. 2026-08-17.(Not yet peer reviewed; the paper shared with outside experts)
- Clay Mathematics Institute. Riemann Hypothesis. Accessed 2026-10-03.(Statement of the hypothesis and the first ten trillion zeros)
- Lean FRO. Lean: Programming Language and Theorem Prover. Accessed 2026-10-03.(Proof checking by a minimal trusted kernel)
- Clay Mathematics Institute. Rules for the Millennium Prize Problems. Accessed 2026-10-03.(Two years after publication and general acceptance as prize conditions)
