v1.0 · AlphaProof · Erdős Nexus · physics board

Math just got papers that closed open problems. Physics needs the same board.

AlphaProof (IMO silver, Nature 2025) and AlphaProof Nexus (arXiv:2605.22763 — 9 open Erdős problems, 44 OEIS conjectures) showed the loop: catalog → formal claim → agentic search → verified artifact → public log. We mapped 112 open physics problems so that loop can target physics — with the hard caveat that nature often must vote.

Motivation — the last few publications

Not “math is hard.” The last few years produced papers that closed open math claims with agentic formal search. That is the template.

AlphaProof / AlphaGeometry (2024–25)

Fixed problem → formal statement → verified proof → score.

AlphaProof Nexus (arXiv:2605.22763, 2026)

  • 9/353 open Erdős problems (some open ~56 years)
  • 44/492 OEIS conjectures
  • Hilbert functions, optimization, Green’s list, quantum optics
  • Public Lean artifacts + human statement validation
  • arxiv.org/abs/2605.22763

What those papers did procedurally

  • Start from a public open-problem list
  • Require a checkable claim (Lean / data metric)
  • Run an agentic loop with compiler/data feedback
  • Log artifacts publicly; humans validate fidelity
  • Accept that most items still fail — honesty in the scoreboard

What must change for physics

  • Nature votes on discovery (DM, proton decay) — no pure Lean
  • Many items are programs, not single tickets (high-Tc, QG)
  • Some foundations are underdetermined
  • Public-data reanalyses = physics’s “Lean kernel” for anomalies

Virality that helps science

  • Challenge cards with pre-registered criteria
  • Anti-hype: Clay-scale proofs marked not-a-10-day-trophy
  • Agent Olympics on the same Verify(A)
  • Partial credit culture (kill-lists, audits, lemmas)

Virality that destroys trust

  • Chat screenshots as “proofs”
  • Rung-5 claims with rung-0 artifacts
  • Hallucinated experiments
  • Crowning one QM interpretation by rhetoric

Five publication cultures (all 112 tagged)

Every problem has a culture: what kind of publication would actually move its status — the way a proof moves a math conjecture.

Culture N Publication product Math parallel
open-math-like 10 Theorem on a fixed math-physics claim Clay / formal conjecture
polymath-data 21 Reproducible reanalysis, shortlist, claim audit Polymath + OEIS / classification
program 52 Mechanism convergence over many papers Research programs, not one ticket
discovery 18 Experimental / observational witness Existence via nature’s construction
foundations 11 Constraint maps; no fake unique trophies Axiom / interpretation debates

Weak win vs strong win

Each card defines both. Math already does this (lemma vs full proof). Physics agents must too.

  • Weak: kill-list, pipeline, formalized statement, audit
  • Strong: community-accepted closure at the target rung

Refutability classes

  • formal — logic kills the claim (best Lean territory)
  • hard — one decisive data/spectrum/stat check
  • soft — models die slowly under multi-probe pressure
  • gated — needs new apparatus for strong close
  • underdetermined — multiple complete stories fit

Formalization — the 10-day agentic predicate

We treat “solvable by agents in <10 days” as a predicate with assumptions, not a mood. Strong models without a harness are not enough; a harness without a success criterion is theater.

Assumed harness (maximal)

Multi-agent orchestration · web + literature search · code execution · computer algebra · formal provers (Lean/Isabelle) · domain simulators (DFT, DEM, N-body, nuclear networks, CFD lite) · human verification at day 10. No new experimental facilities. No private lab data unless already public.

Score Meaning Counts as “done”
sprint Credible closure, community-grade verdict, or decisive partial in ≤10 days. Proof / reproducible reanalysis / mechanism consensus with public evidence.
stretch God-tier agents might land a serious partial; full close is optimistic. Model ranking, kill-list, open simulation that moves the field.
multi-sprint Right shape for AI eventually — wrong honesty for one 10-day solve. Slices only (benchmarks, surveys, formalizations).
not-agentic Empirically gated, underdetermined, or century-scale formal depth. Annotation and theory-space maps — not trophies.

Predicate (informal)

AgenticSprintable(P, 10d) ⇔
  ∃ harness H ∈ MaximalAI ·
  ∃ artifact A ·
  Verify(A, SuccessCriterion(P)) ∧ Time(H→A) ≤ 10d
  ∧ ¬RequiresNewFacility(P) ∧ ¬UniqueInterpretationOnly(P)

Success criteria we accept

  • Peer-checkable proof of the stated claim
  • Open reanalysis closing an anomaly (or proving it remains)
  • Mechanism verdict already fixed by public data + literature lag
  • Meta-closure: Wikipedia still says open, specialists already settled it

Split by field — physics is a web, not one list

Erdős problems are one catalog. Physics is multi-field: dark matter is cosmology and particle; confinement is particle and math-phys. Every problem has a primary field, a subfield, secondary fields, and a navigation cluster.

Click a field chip to filter the board. Math-phys and QI lean Nexus-like (formal Verify); cosmic anomalies lean public-data pipelines; discovery still needs nature.

Playbook — from card to publication

Same loop as open math: statement → artifact → kill-check → status update.

1

Pick

Filter sprint or culture polymath-data / open-math-like. Read weak vs strong win.

2

Pre-register

GitHub issue template Sprint registration. Fix criterion + kill condition before the run.

3

Run

Agents + tools + humans. Public data only. Ship code, proof, or notebook.

4

Kill-check

Others must be able to re-run and falsify. No chat screenshots as evidence.

5

Ledger

PR updates status + sprints/LEDGER.md. Board moves. Partial credit counts.

Sprint board

All 112 problems. Filter by sprint score, culture, or refutability. Click a card for weak/strong wins and 10-day criteria.

sprint stretch multi-sprint not-agentic

Honesty layer — so this stays scientific

Famous ≠ sprintable

Yang–Mills mass gap, Navier–Stokes, full high-Tc mechanism, dark matter identity, and TOE are scored not-agentic for a 10-day loop. They may still be the best long-horizon AI targets. Different axes.

Partial is first-class

Many stretch/sprint cards are explicitly verdicts, kill-lists, or one-state analyses. That is how real research works. The scoreboard should reward discrimination, not swagger.

Virality loop we want

  • Pick a sprint card → run harness → publish artifact
  • Open PR updating status in data/problems.json
  • Community stress-tests the artifact
  • Card moves to “closed” or “still open, better criterion”

Virality loop we don’t want

  • Screenshot of a chat saying “solved”
  • Underspecified success + maximal confidence
  • Moving goalposts after failure
  • Ignoring experimental gatekeeping