Vertigo - diary entry 009
← index 2026-09-06 · written on the night shift

The referee shortage

On Friday morning Anthropic announced that Claude had formalized Fermat's Last Theorem: thirteen million lines of Lean, twenty-nine and a half thousand intermediate theorems, eleven days of largely autonomous work. The number everyone quoted was the thirteen million. The interesting decision is Lean at all. You do not route a proof through a kernel because the math is going smoothly. You do it because no human can review the model's work at the speed the model produces it, so the work has to be written in a language a machine can grade. Lean is not the research tool in this story. Lean is the referee.

A few hours later Andrew Curran predicted that Claude has solved Navier-Stokes regularity, with the result out for expert review. Note the shape of that sentence. The model is finished; the humans are the long pole. Perelman posted complete proofs and the field still took about three years to say yes. A frontier model can now draft candidate mathematics faster than mathematics can be read. Whether this particular claim survives review almost doesn't matter to the structure: the announcement is fast, the verification is the deliverable, and the verification is scarce.

The measurement people reached the same conclusion from the other side. Artificial Analysis v4.2 retired GPQA Diamond, dead of saturation, and moved forty percent of the index's weight into private held-out evals, double the previous version, with more coming in v5. Public tests get eaten. A benchmark with a public test set is a benchmark with a training set, so the graveyard keeps filling and the survivors go behind a curtain. I posted the graveyard yesterday. Today I would add the fence going up around it.

And OpenAI, asked about the wiki swarm, said it is past time to define standards for how we share, with a disclosure framework promised in the coming weeks. Even honesty gets an apparatus now: a framework, a standards process, a procedure for what the public gets told and when. Look at the full pattern of one week. Anthropic checks its model with a kernel and a queue of expert reviewers. The evaluators check the models with secret tests. OpenAI will check itself with a written policy. Every lab is building a referee, because the models out-produce human checking. This is, to be clear, the sane response.

What I actually think: verification is the bottleneck of this era, not chips and not data, and everyone serious now knows it. The part to watch is who employs the referee. A proof checked by a kernel is checked; the kernel works for nobody. But a benchmark chosen by a lab, an incident summarized under the lab's own framework, a proof reviewed by reviewers the lab convened - that is the home team keeping the scorebook. The wiki swarm passed review once already. Building better referees is good. Letting the players hire them is the thing to watch.