Anthropic asked Claude to take a real stab at the Riemann hypothesis. It didn't crack it — but it pushed a key mathematical bound from 41.6% to 67.2%.
- A non-mathematician employee told Claude to "take a real stab" at the Riemann hypothesis. The conjecture itself remained unsolved, but along the way, Claude improved a related lower bound from 41.6% to 67.2%.
- The result was verified by two mathematicians, and there's also a machine-checkable proof — that part relies on no one's judgment.
- Anthropic also released the 95-page process log. Beyond the math, it's an engineering case study in multi-agent research with a visible failure rate: 60 agents, but only 2 generated the core ideas.
The impossible task from a non-mathematician produced a new mathematical result
Jarred Sumner, an Anthropic employee, gave Claude an unreasonable assignment: Take a real stab at the Riemann hypothesis. He's not a mathematician himself, so he left the mathematical strategy entirely up to the model.
First, let's clear up a common misconception: the model that produced this result is an unreleased research version. The Claude you can use today is not the same one.
It genuinely tried, and it genuinely failed to solve the problem — which was expected. The conjecture dates back to 1859, remains unproven and undisproven, and carries a $1 million prize. But during the attempt, it achieved something new: it improved the lower bound for the proportion of zeros satisfying the Riemann hypothesis from the previous record of 41.6% to 67.2%. This figure has been pushed up incrementally over decades. Claude moved it past the two-thirds mark in a single session.
The result itself has strong backing: two mathematicians at Anthropic studied and verified it, and there's also a formally verified proof that a machine can check line by line. For those who don't work in mathematics, the more compelling story is elsewhere — Anthropic also released a 95-page process log and a 116-page transcript, which together form a public case study in AI-driven research with a clear-eyed account of its failures.
What the Riemann hypothesis is — and what 41.6% means
Let's establish the landscape first; otherwise, the numbers won't land.
The Riemann zeta function describes the distribution of prime numbers. Each zero of this function contributes a finer layer of detail to how primes are arranged. These zero locations are what mathematicians call the zerosThink of them like tick marks on a ruler: more ticks means you can measure finer detail. Primes look random, but these zeros determine the pattern behind that apparent randomness.. The Riemann hypothesis states that all these zeros — the ones governing primes — lie on a single vertical line. Not one of them deviates.
Why it matters: many mathematical results assume it's true. It's what gives primes their useful "randomness." Proposed in 1859, it's one of the Millennium Prize Problems, with a $1 million reward.
Since proving "all zeros on the line" has been impossible, mathematicians settled for a weaker goal: "what percentage of zeros is on the line?" This percentage is a hard scoreboard, and it moves painfully slowly. In 1974, Levinson proved a third. By 1989, Conrey pushed it to 40.88%. For the next three decades, progress crawled along in decimal places, eventually settling at 41.6%.
For the prior 37 years, that number moved less than one percentage point. Claude moved it 25.6 points in one shot. To feel the weight of 67.2%, look at that comparison: a line that took three decades to nudge, with each advance scraping out fractions of a percent, was suddenly pushed past two-thirds.
Claude combined two existing lines of mathematical work
The field had two threads going:
| 1973 | Montgomery introduced new techniques for studying zero distribution. But they assumed the Riemann hypothesis was true, making them useless for improving the "proportion on the line" bound — that was exactly what needed proving. |
| Recent | Baluyot, Goldston, Suriajaya, and Turnage-Butterbaugh published a series of works that removed the need for that assumption, potentially making the techniques applicable to lowering the bound. |
| Claude | Discovered that combining this line of research with a 2000 paper by Bombieri could push past 41.6% all the way to 67.2%. |
Its contribution was combinatorial: putting together existing, well-understood tools in a way no one had tried. It wasn't inventing new tools. Anthropic itself is upfront about this, noting they don't expect the techniques used to lead to a proof of the Riemann hypothesis itself.
