Mathematics feels on the edge of a crisis, and the culprit is surprisingly a result an AI model obtained on a problem connected to the Riemann hypothesis. Anthropic's unreleased research version of Claude raised the decade-old lower bound on the proportion of zeta-function zeros on the critical line from 41.6% to 67.25%. Ellie Sleightholm, a mathematics PhD student at Cambridge, takes up this development in the first episode of her new series on mathematics in the age of AI, because almost everything circulating online as the hypothesis is proved or the hypothesis has fallen is wrong.
It all starts with the atoms of numbers: prime numbers , integers above 1 divisible only by 1 and themselves. Every integer factors uniquely into primes, as in 60 = 2 × 2 × 3 × 5. The prime-counting function π(x) tells how many primes sit below x; π(10) is four, since the primes below 10 are 2, 3, 5 and 7. Its graph climbs in steps: flat stretches with no new prime, sudden jumps at each prime. Gauss and Legendre conjectured at the end of the 18th century that this staircase hugs x divided by log x, and the prime number theorem was proved in 1896 by Hadamard and de la Vallée Poussin. We know the average; what we do not know is how wildly the deviations from that average can swing.
From primes to the zeta function
In the 1730s Euler found that the sum of reciprocal squares, 1 + 1/4 + 1/9 + ..., converges to π²/6; this is the Basel problem. Replacing the 2 with s gives 1/1ˢ + 1/2ˢ + 1/3ˢ + ..., the Riemann zeta function , and setting s = 2 recovers the Basel sum. The geometric-series formula turns, with R = p⁻ˢ, into one factor per prime, and taking Euler's infinite product over all primes yields the Euler product of the zeta function. Its expansion produces every integer's reciprocal exactly once, thanks to the fundamental theorem of arithmetic, as in 12 appearing as 1/(2²×3). The function describing the distribution of primes is thus multiplicatively wired to the primes themselves.
In 1859 Riemann published a paper of a few pages, on the number of primes less than a given magnitude, asking a brilliant question: what if Euler's definition is extended to complex numbers? Writing s = σ + it lets s roam the whole complex plane, as in plugging in 3 + i instead of 3. Where the real part exceeds 1 the series converges neatly and produces mesmerizing transformations. Representing the remaining half-plane uses analytic continuation , the unique extension that complex analysis permits. The same extension explains why 1 + 2 + 3 + ... is associated with −1/12. In this form the zeta function directly constrains the growth of the prime count, that is, how faithfully primes follow the x/log x curve.
The inputs that make the zeta function vanish, the complex numbers with ζ(s) = 0, are the heart of the story. The trivial zeros at −2, −4, −6 and the other negative even integers come from a sine factor in the formula and excite nobody. What matters are the nontrivial zeros ; Riemann showed they live in the critical strip where the real part of s lies between 0 and 1. The hypothesis claims every one of them sits on the critical line with real part exactly 1/2. This is the problem on ClayMath's (the Clay Mathematics Institute) Millennium list with a one-million-dollar prize: the prime number theorem governs the average distribution, the hypothesis governs the deviation from it. The first nontrivial zero sits near 1/2 + 14.134725i, and a single counterexample, one zero off the line, would destroy the hypothesis.
From 1859 to 2026: a history of proportions
The proportion idea has grown step by step for a century. In 1914 Hardy proved infinitely many zeros lie on the critical line, intuitively beautiful yet no closer to the hypothesis, since infinitely many could still lie off it. In 1942 the Fields medalist Selberg showed a positive proportion of the zeros sits on the line, locking in a fixed percentage. On the speaker's chronology, Levinson lifted it to 34.7% in 1974, Conrey to 40.9% in 1989, and 41.7% was seen in 2022. As Conrey's survey in the AMS Notices recalls, Hilbert placed the problem among his 23 problems in the 1900 Paris address; the curiosity of that day has become today's industry. On the refereed track, Bui, Conrey and Young crossed the 41% bar in 2011, and the official record Anthropic inherited was 41.6%.
Anthropic announced the result on August 10, 2026 with an X post and an explanatory essay; as TechSpot reported, the experiment ran on an unreleased research version. The first sentence of the message matters most: Jarred Sumner, a non-mathematician on staff, asked Claude to take a real stab at the hypothesis itself, the model tried and failed, but made unexpected progress on the related problem along the way. That is the knot of the social-media confusion: while the company said it did not solve it but advanced the neighboring question, posts multiplied claiming it is proved, it has fallen, it will fall next year. Confronting that flood of misinformation is the main motive of Sleightholm's video.
Sixty agents, a day and a half
The scale of the orchestration is startling. In the first round Claude generated and tried 650 ideas and none worked, a picture every mathematician recognizes. In the second round Sumner pushed the model again, and Claude coordinated about 60 subagents for a day and a half. Per Anthropic's footnote, two of these agents developed the core mathematical ideas, 13 fed ideas to those two, 30 tried new ideas without success, 13 served as checking referees, and the last two wrote the first paper draft. Together they ran 2,400 shell commands, wrote hundreds of Python scripts and spent 31 million output tokens, with thousands of numerical checks against known zeta zeros. Sumner's own input was mostly encouragement, variations on keep going and believe in yourself, which seems to have helped break the model's initial skepticism.
The turning point grew out of a failure. Claude arranged the problem as a six-rung ladder and on the top rung tried to build the kind of mathematical object that would force the hypothesis to be true; it failed. But an agent called R4 noticed the structure under study did not behave like an ordinary positive inner-product space: it had positive and negative directions, an indefinite structure. First counted as an obstacle, this became an asset about 12 hours later: another agent, E2, flipped the argument and instead of using the negative part to count off-line zeros tried pinning on-line zeros with the positive part. From this came the claim that at least half the zeros are distinct and on the critical line. Claude stayed skeptical and sent in independent referee agents, deliberately blinded to each other; they attacked the localization, the prime side and the linear algebra, genuinely found a matrix gap, and more importantly repaired it. The 50% threshold held.
What followed is a story of fine-tuning and stubbornness. Refining the mathematical window moved the bound only from 50% to about 50.66%, and a higher-moments attempt hit a theoretical ceiling. Remaining hope rested on a factor called E2 pairs; the agent wrote the proof to disk and then crashed on an infrastructure fault, so the coordinator inspected the abandoned proof line by line, restarted the agent and sent the result to two breaker agents at once. The outcome: at least two thirds of the zeros are distinct and on the critical line, polished at the end to 67.25%. The Alpöge-Furman paper on arXiv formalizes the result: at least two thirds in weighted count, at least five sixths of the distinct zeros, with constants 0.6725 and 0.8362 under the Montgomery-Taylor window. It makes Montgomery's 1973 deduction unconditional by combining a finite-dimensional matrix picture of the Weil Hermitian form with Sylvester's law of inertia , using no mollifier, zero-density estimate or zero-free region.
Verification, and why it is no proof
To secure the result, Claude had subagents review the proofs, hunt for counterexamples, re-derive the finding independently from scratch, and download 54 papers from the arXiv to check the result was not already known. It then volunteered to write the finding as a paper and insisted the next reader should be human, requiring validation by a number theorist. Anthropic's own mathematicians, Levent Alpöge and Ralph Furman, examined and verified the work, while Brian Conrey and Dan Goldston looked at the paper on short notice. In parallel a formal proof was produced in Lean with Eric Easley; Lean is a proof assistant in which a computer checks every logical step, and the result passed the standard validator. The chain shows the minimum seriousness a machine-made mathematics claim needs before it is taken seriously.
And the most critical point: the result counts as evidence for the hypothesis in neither direction; in the coordinator's own words, the on-line proportion offers no evidence about the truth of the hypothesis either way. Even reaching 100% would change nothing, and Anthropic states that technical limit openly. Sleightholm builds the intuition with perfect squares: among the first n integers about n − √n are not squares, a proportion 1 − 1/√n tending to one, so even almost 100% not-a-square still leaves infinitely many squares. Off-line zeros could likewise be sparse enough to vanish in the total count yet still infinite. So 67.25% is an impressive lower bound, not a conclusion.
Sleightholm moves from here to five lessons. First, pace: a few years ago we asked a chatbot single questions; today a team of sixty agents talks, divides labor and works independently, like a research group at a university. Second, guardrails: imagining a malicious actor with access to an agent team of this scale is chilling; she praises Anthropic for publishing misuse research and wants pressure on the companies. Third, Millennium problems turned into AI benchmarks: systems generate hundreds of pages of proofs while the mathematics community lacks referee capacity to digest them; OpenAI pulling the credits it offered for a math marathon after the backlash shows the two communities must work together. Fourth, transparency: Anthropic opened the logs and said we tried, we failed, here is what we found, whereas OpenAI's Navier-Stokes announcement initially omitted the two mathematicians whose groundwork enabled it. Fifth is Tao's question: mathematics was never only proof production; new ideas, understanding why results hold, explaining, connecting fields and passing knowledge on will matter even more in an era of proof abundance.
She closes in her PhD-student voice: watching her whole field shift, she feels excited and frightened at once, and invites viewers to debate that future together. The video opens with a mathematics-in-crisis joke and ends with a community call; the 35 minutes in between pack a complete mini-course from primes to the zeta function and from the Euler product to the critical line, topped with the backstage story of the agent orchestration. Viewers walk away knowing both what the hypothesis says and what a language model with a sixty-strong invisible team can do in a day and a half.
Key moments
- Opening: a sense of crisis in mathematics and Riemann-linked progress
- Series launch: mathematics in the age of AI
- Primes and the prime-counting function
- From the Basel problem to the zeta function
- Riemann's 1859 paper and analytic continuation
- The critical line and the statement of the hypothesis
- History of proportions: Hardy to Selberg
- 650 failed ideas and the 60-agent second attempt
- How the subagents divided the work
- The 50% claim under blind referee scrutiny
- A 54-paper literature sweep and validation
- Formal proof in Lean
- The square-number analogy: why a proportion is no proof
- Tao on the proof-abundance era and closing
AI commentary
"This result both excites and unsettles me: a sixty-agent team broke a decade-old number-theory bar in a day and a half. Yet readers should look at the proof itself, not the headlines."
AI assessment
The strongest objection comes from the numbers themselves: however high the on-line proportion climbs, it counts as evidence for the hypothesis in neither direction, and the person saying so is Claude's own coordinator. The Lamzouri paper on arXiv reaches the same 67.25% and 83.62% figures with a shorter Hilbert-space inequality and adds two new estimates, so the conceptual load still rests on human mathematics: Montgomery's 1973 techniques, the unconditionalization work of Aryan and of Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh, and Bombieri's 2000 paper. Anthropic's own essay states plainly that the result builds on decades of that accumulated work, which makes the solved-it-single-handedly narrative an exaggeration.
The limits are equally open: the paper is a preprint without peer review, and the Conrey-Goldston look was a short-notice examination. The Lean formalization built with Easley shows the statement passes machine checking, so the chain of logic is computer-verified, but it grades neither the originality nor the importance of the proof. On the plus side, Anthropic opened the process logs, the failed 650 ideas and the agent logs; unlike comparable claims, that transparency is what makes the claim worth taking seriously.
The speaker's position deserves a note too: Sleightholm is a Cambridge PhD student launching a series on mathematics in the age of AI, and the subscription plea plus community-building motive explain the excited tone. In the video she shares a personal discount code (Ellie 10) for the wool-clothing brand Outside In; by her own account not a paid promotion and tied to a donation model for the homeless, yet personal brand and content share the same frame. That she herself voices mathematicians' general transparency criticism of AI companies honestly lays out this balance of interests.
The practical takeaway for readers: when you see headlines that Riemann is proved or the hypothesis has fallen, ask what the 67.25% actually counts, the proportion of simple zeros on the critical line, not the hypothesis itself. Those who want to follow the subject seriously should read three documents: the Alpöge-Furman preprint on arXiv, Anthropic's explanatory essay and the Lean record. For young researchers the real lesson is methodological: a skeptical coordinator, blinded checkers that never see each other's work, a 54-paper literature sweep, and the next-reader-is-human principle, a discipline anyone working with agents can copy.
Sources
6 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.
riemann hypothesis · claude · anthropic · prime numbers · zeta function · ai and mathematics · lean