Math just got papers that closed open problems. Physics needs the same board.
AlphaProof (IMO silver, Nature 2025) and AlphaProof Nexus (arXiv:2605.22763 — 9 open Erdős problems, 44 OEIS conjectures) showed the loop: catalog → formal claim → agentic search → verified artifact → public log. We mapped 112 open physics problems so that loop can target physics — with the hard caveat that nature often must vote.
Motivation — the last few publications
Not “math is hard.” The last few years produced papers that closed open math claims with agentic formal search. That is the template.
AlphaProof / AlphaGeometry (2024–25)
- IMO 2024: 4/6 problems, silver-medal level
- Formal proofs in Lean — kernel-checked, not chat
- Nature 2025 methodology paper
Fixed problem → formal statement → verified proof → score.
AlphaProof Nexus (arXiv:2605.22763, 2026)
- 9/353 open Erdős problems (some open ~56 years)
- 44/492 OEIS conjectures
- Hilbert functions, optimization, Green’s list, quantum optics
- Public Lean artifacts + human statement validation
- arxiv.org/abs/2605.22763
What those papers did procedurally
- Start from a public open-problem list
- Require a checkable claim (Lean / data metric)
- Run an agentic loop with compiler/data feedback
- Log artifacts publicly; humans validate fidelity
- Accept that most items still fail — honesty in the scoreboard
What must change for physics
- Nature votes on discovery (DM, proton decay) — no pure Lean
- Many items are programs, not single tickets (high-Tc, QG)
- Some foundations are underdetermined
- Public-data reanalyses = physics’s “Lean kernel” for anomalies
Virality that helps science
- Challenge cards with pre-registered criteria
- Anti-hype: Clay-scale proofs marked not-a-10-day-trophy
- Agent Olympics on the same Verify(A)
- Partial credit culture (kill-lists, audits, lemmas)
Virality that destroys trust
- Chat screenshots as “proofs”
- Rung-5 claims with rung-0 artifacts
- Hallucinated experiments
- Crowning one QM interpretation by rhetoric
Five publication cultures (all 112 tagged)
Every problem has a culture: what kind of publication would actually move its status — the way a proof moves a math conjecture.
| Culture | N | Publication product | Math parallel |
|---|---|---|---|
| open-math-like | 10 | Theorem on a fixed math-physics claim | Clay / formal conjecture |
| polymath-data | 21 | Reproducible reanalysis, shortlist, claim audit | Polymath + OEIS / classification |
| program | 52 | Mechanism convergence over many papers | Research programs, not one ticket |
| discovery | 18 | Experimental / observational witness | Existence via nature’s construction |
| foundations | 11 | Constraint maps; no fake unique trophies | Axiom / interpretation debates |
Weak win vs strong win
Each card defines both. Math already does this (lemma vs full proof). Physics agents must too.
- Weak: kill-list, pipeline, formalized statement, audit
- Strong: community-accepted closure at the target rung
Refutability classes
- formal — logic kills the claim (best Lean territory)
- hard — one decisive data/spectrum/stat check
- soft — models die slowly under multi-probe pressure
- gated — needs new apparatus for strong close
- underdetermined — multiple complete stories fit
Formalization — the 10-day agentic predicate
We treat “solvable by agents in <10 days” as a predicate with assumptions, not a mood. Strong models without a harness are not enough; a harness without a success criterion is theater.
Assumed harness (maximal)
Multi-agent orchestration · web + literature search · code execution · computer algebra · formal provers (Lean/Isabelle) · domain simulators (DFT, DEM, N-body, nuclear networks, CFD lite) · human verification at day 10. No new experimental facilities. No private lab data unless already public.
| Score | Meaning | Counts as “done” |
|---|---|---|
| sprint | Credible closure, community-grade verdict, or decisive partial in ≤10 days. | Proof / reproducible reanalysis / mechanism consensus with public evidence. |
| stretch | God-tier agents might land a serious partial; full close is optimistic. | Model ranking, kill-list, open simulation that moves the field. |
| multi-sprint | Right shape for AI eventually — wrong honesty for one 10-day solve. | Slices only (benchmarks, surveys, formalizations). |
| not-agentic | Empirically gated, underdetermined, or century-scale formal depth. | Annotation and theory-space maps — not trophies. |
Predicate (informal)
AgenticSprintable(P, 10d) ⇔
∃ harness H ∈ MaximalAI ·
∃ artifact A ·
Verify(A, SuccessCriterion(P)) ∧ Time(H→A) ≤ 10d
∧ ¬RequiresNewFacility(P) ∧ ¬UniqueInterpretationOnly(P)
Success criteria we accept
- Peer-checkable proof of the stated claim
- Open reanalysis closing an anomaly (or proving it remains)
- Mechanism verdict already fixed by public data + literature lag
- Meta-closure: Wikipedia still says open, specialists already settled it
Split by field — physics is a web, not one list
Erdős problems are one catalog. Physics is multi-field: dark matter is cosmology and particle; confinement is particle and math-phys. Every problem has a primary field, a subfield, secondary fields, and a navigation cluster.
Click a field chip to filter the board. Math-phys and QI lean Nexus-like (formal Verify); cosmic anomalies lean public-data pipelines; discovery still needs nature.
Playbook — from card to publication
Same loop as open math: statement → artifact → kill-check → status update.
Pick
Filter sprint or culture polymath-data / open-math-like. Read weak vs strong win.
Pre-register
GitHub issue template Sprint registration. Fix criterion + kill condition before the run.
Run
Agents + tools + humans. Public data only. Ship code, proof, or notebook.
Kill-check
Others must be able to re-run and falsify. No chat screenshots as evidence.
Ledger
PR updates status + sprints/LEDGER.md. Board moves. Partial credit counts.
Sprint board
All 112 problems. Filter by sprint score, culture, or refutability. Click a card for weak/strong wins and 10-day criteria.
Honesty layer — so this stays scientific
Famous ≠ sprintable
Yang–Mills mass gap, Navier–Stokes, full high-Tc mechanism, dark matter identity, and TOE are scored not-agentic for a 10-day loop. They may still be the best long-horizon AI targets. Different axes.
Partial is first-class
Many stretch/sprint cards are explicitly verdicts, kill-lists, or one-state analyses. That is how real research works. The scoreboard should reward discrimination, not swagger.
Virality loop we want
- Pick a sprint card → run harness → publish artifact
- Open PR updating status in
data/problems.json - Community stress-tests the artifact
- Card moves to “closed” or “still open, better criterion”
Virality loop we don’t want
- Screenshot of a chat saying “solved”
- Underspecified success + maximal confidence
- Moving goalposts after failure
- Ignoring experimental gatekeeping