Anthropic published a guest research post from physicist and science writer Matt von Hippel: Claude, running Fable 5.1 inside the Claude Science harness, computed the six-particle (hexagon) scattering amplitude in planar N=4 super Yang-Mills at nine loops — one loop past the eight-loop record SLAC’s Lance Dixon published in 2023.
Cost per method: roughly $1,000–$2,000 of Claude Science usage, with about $100 of that going to the SymPy bootstrap on the equivalent of 96 CPUs for a week. Supervision was thin: a short problem statement, then “keep working; update every 4–6 hours.” Dixon independently validated the result. A Chinese Academy of Sciences group led by Song He reached most of the same nine-loop symbol around the same time with GPT-6 assistance inside a human-built framework.
Image credit: Anthropic
Key points
- What shipped: An Anthropic research post + public nine-loop amplitude/form-factor data (Cosmic9 / Zenodo), not a new consumer model.
- Model + harness: Claude Fable 5.1 inside Claude Science — a paid science harness with structured rules/prompts.
- Task: Six-particle MHV amplitude in planar N=4 SYM at nine loops (Dixon’s prior human record: eight).
- Methods: Claude ran two routes — direct bootstrap and the indirect form-factor / antipodal-duality path Dixon used for eight loops.
- Cost: ~$1–2k per method for the end user; ~$100 of bootstrap CPU time (96-CPU-week class).
- Validation: Lance Dixon (SLAC / Stanford) checked the result; Cosmic9 publishes symbols, samples, and validation records.
- Parallel work: Song He, Jirong Jing, and Xiang Li (CAS) released a concurrent nine-loop symbol dataset with GPT-6 help — less “one-shot autonomous” than Anthropic’s framing.
- Honest limit: Claude executed known recipes with more compute and better software engineering than many academic setups — it did not invent a new physical principle. von Hippel and Dixon both say the next soul-searching moment is novel insight, not recipe execution.
What shipped
Von Hippel had posted a public challenge: show that an AI can take academic-scale compute and solve a frontier amplitudes problem — specifically N=4 SYM to nine loops (or N=8 supergravity to seven). Anthropic physicists Liam Fitzpatrick and Siddharth Mishra-Sharma took the N=4 challenge.
They asked Claude which problem it was most confident about, then gave a short prompt along the lines of: compute the six-particle (hexagon) amplitude in planar N=4 SYM at nine loops. After that, the standing instruction was keep going through off-hours with periodic updates.
Claude finished twice — bootstrap and form-factor routes — each in the low-thousands of dollars. The bootstrap used Python + SymPy. Results are posted in the Cosmic9 layout Dixon’s group already uses for lower loops, with checksums, sample coefficients, and method/validation notes. Large files sit on Zenodo.
Anthropic invited von Hippel to write the guest post and paid him; Dixon validated independently and received Claude usage credits. The post discloses both.
What changed
Amplitudes are hard because each extra “loop” multiplies complexity. Most real-world scattering formulas stop at two or three loops; Dixon’s eight-loop toy-model result was already a landmark. The bootstrap method — constrain a symbolic alphabet until only one answer survives — is fragile: one recipe bug collapses the whole soufflé.
Three things matter beyond the headline “AI did physics”:
- Harness reliability over a week. von Hippel’s takeaway is that Claude Science finished a finicky, multi-day computation with almost no scientific babysitting. Humans often burn a second week debugging the first week. That is a product reliability story as much as a science story.
- Cost accessibility. A frontier toy-model amplitude at individual-researcher price — not national-lab scale — changes who can attempt the next loop.
- Simultaneous convergence. CAS + GPT-6 hitting the same target in the same window undercuts “only Anthropic’s secret sauce” narratives. Humans still get to publish, explain, and analyze; Claude’s role in this episode is computation + packaging.
Dixon’s addendum is unusually clear-eyed: Claude understands the 2019 and 2023 papers “better than any human aside from my co-authors,” used their methods, and presented output in their format — so validating Claude also re-validates years of human work. He is not crushed; he is waiting for models that invent new principles.
What it means for builders and scientists
1. Treat Claude Science as the product surface, not “Fable did math.”
If you run long research agents, the interesting stack is harness + standing “keep going” ops + dual-method cross-checks + an expert who can invalidate the soufflé. The model alone is not the story.
2. Budget for autonomy as a reliability metric.
A week-long unsupervised run that survives fragile symbolic pipelines is a different bar than chat demos. Ask vendors (and your own agents) for: dual independent methods, sample-coefficient checks, and a human oracle who knows the domain.
3. Do not overclaim discovery.
This is recipe execution at frontier loop order with known techniques. Useful, impressive, and builder-relevant — still short of “AI found a new law of nature.” Keep that distinction in decks and papers.
4. Watch real-world amplitudes next.
Toy models (N=4) are friendlier than LHC-facing calculations. von Hippel expects more low-hanging fruit but warns the competitive, multi-group real-world lane may have less. Builders in scientific computing should pressure-test harnesses on problems with independent validation paths.
5. Attribution and publishing norms are still human-paced.
Dixon and Song He’s groups publish and explain; Claude’s run stops at the data release. Labs using agents for science need clear credit rules before simultaneous AI + human convergence becomes the default.
What to watch
- Coefficient-level match between Cosmic9 / Claude data and the full CAS GPT-6-assisted release.
- Whether Claude Science (or rivals) one-shot a real-world amplitude loop that groups thought was out of reach.
- How Anthropic prices and documents long unsupervised science runs for outside researchers.
- The first case where an LLM proposes a new amplitude technique humans then adopt — Dixon’s stated bar for a deeper shift.



