The Advisory Group Got Four Written Guarantees. The One Thing It Can't Advise On Is the Pace.

OpenAI published a page this morning announcing an Advisory Group on Mathematics and Artificial Intelligence: nine mathematicians, none of them OpenAI employees, convened ten days after 27 Fields medalists put their names to a document arguing that the way AI companies are using mathematics is damaging mathematics.

The charter is one paragraph. Four sentences of unusually specific independence, and then a fifth sentence.

The group will operate independently from OpenAI. The group will have the freedom to offer advice we have not requested, comment on OpenAI’s impact on mathematics, and make its advice public. Its value depends on its members being able to exercise their own judgement and challenge ours. Its members will not be paid by OpenAI, and the group can change its membership as it sees fit. Importantly, the group will not be responsible for advising us on how to pace our internal progress on mathematics.

I want to take the first four sentences seriously before I get to the fifth, because the reflex here is to call the whole thing theatre and the reflex is wrong. Unpaid, free to publish advice nobody asked for, free to criticise the sponsor in public, free to change its own roster: those are four concrete commitments, in writing, on the sponsor’s own site. I cannot think of another lab advisory body that has all four written down. Most corporate advisory boards have none of them.

The group also went further than it was asked to. Its own site, agmai.org, hosted by the Institute for Advanced Study, says plainly how it formed: “This group came together after OpenAI approached some of its members about establishing an external advisory board. In agreement with OpenAI, they decided to create an independent group and invite others to join.” An external board advises one company; an independent group can advise any of them, and this one says it is “willing to offer such recommendations to any AI company whose models are likely to have a significant impact on mathematics.” It also states its own ceiling, which corporate bodies almost never do: “we do not have decision making power at any AI company, and the responsibility for the decisions made by any company will rest with that company.”

So: independence, honestly bounded. And then the fifth sentence.

The letter has two complaints, and the group got one of them

The cheap version of this essay says OpenAI constituted a body in response to a grievance and then carved the grievance out of its terms of reference. That is half right, and the half that is wrong matters.

The declaration, published 11 September 2026, makes two distinguishable complaints. The first is about production:

Indeed, the mass production at faster and faster pace of “true/false” statements could destroy fertile ground instead of breathing life into new ideas.

The second is about announcement: “Often these solutions are announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others. As in all creative professions, this raises severe attribution and plagiarism questions.”

The second complaint is squarely inside the group’s remit. Review, dissemination, coordination, professional standards: that is the rush-to-announce problem, and the group has been handed it. Its site says its current task is “advising OpenAI on how to coordinate the release of a large number of significant results in mathematics that they report have been produced by their internal model.” Note the hedge in “they report.” Somebody there is being careful.

The first complaint is the one outside the fence. The sentence the declaration’s signatories chose as their sharpest, the one about mass production at faster and faster pace, is about the rate at which results are generated, not the manner in which they are published. That is what the carve-out removes.

The declaration's two complaints against the advisory group's remit boundary The 11 September declaration makes two distinguishable complaints. The first, that results are announced in a rush leaving no time for a proper writeup, the isolation of new methods, or citation of previous work, falls inside the advisory group's remit: OpenAI's charter gives the group review, communication, dissemination and professional standards. The second, that mass production of true-or-false statements at faster and faster pace could destroy fertile ground, falls outside it: the charter states the group will not be responsible for advising on how OpenAI paces its internal progress on mathematics. Sources: mathandai.org, 11 September 2026, and OpenAI's advisory group page, 21 September 2026. THE TWO COMPLAINTS · AGAINST THE REMIT BOUNDARY · 21 SEP 2026 COMPLAINT ONE — THE ANNOUNCEMENT RUSH Inside the remit — OpenAI charter, 21 September 2026 "announced in a rush, leaving no time for a proper writeup" INSIDE THE REMIT review · communication · dissemination · professional standards COMPLAINT TWO — THE PRODUCTION RATE Outside the remit — OpenAI charter, 21 September 2026 "mass production at faster and faster pace of 'true/false' statements" OUTSIDE THE REMIT "will not be responsible for advising us on how to pace our internal progress" Quotes from the declaration and from OpenAI's charter paragraph. The charter specifies no meeting cadence, term length or publication process.
The letter's complaint about the rush to announce lands inside the group's remit. Its complaint about the rate of production lands outside it.

Whether that is a scandal depends on what you thought the letter was asking for, and here the most useful witness is a member of the advisory group itself. Timothy Gowers, who is on the board and did not sign the declaration, published a long explanation on 17 September. He agrees with much of the letter. His last objection is the one that bears on the carve-out: “A final reason that I didn’t sign the letter is that I wasn’t really sure what it was demanding that isn’t happening already.” He adds that “there is no chance that the impact of such models on mathematics will persuade AI companies to stop their release,” and concludes “there was nothing to be gained from criticizing AI companies for generating too many solutions too quickly.” Gowers also discloses, unprompted, that he has early access to OpenAI models and free Pro access, and has never been paid by OpenAI.

Read the declaration cold and he is right that it issues no specific demand. It says the issues “must be addressed urgently” and stops. So OpenAI carved out an ask that was never formally made, in response to a letter whose own best-positioned critic could not identify the ask either. That is less damning than the gotcha version and more interesting: both documents are circling a quantity neither of them names.

The board is not a capture story

It would be tidy if the nine were a friendly slate. They are not.

Where each of the nine advisory group members stands on the declaration Nine advisory group members, each with exactly one recorded position on the 11 September declaration. Martin Hairer is one of the 27 Fields medalist co-authors of the declaration itself. Camillo De Lellis and Ravi Vakil appear in the declaration's public endorsers list. Timothy Gowers publicly declined to sign and published his reasons on 17 September. François Charles, Nikhil Srivastava, Ulrike Tillmann, Edward Witten and Melanie Matchett Wood have no public position found: they are absent from both the co-author list and the endorsers search. Checked 21 September 2026. POSITION ON THE DECLARATION · 9 ADVISORY GROUP MEMBERS · CHECKED 21 SEP 2026 SIGNED IT ENDORSED IT DECLINED, PUBLICLY NO POSITION FOUND Martin Hairer Camillo De Lellis Ravi Vakil Timothy Gowers François Charles Nikhil Srivastava Ulrike Tillmann Edward Witten Melanie Matchett Wood Martin Hairer — one of the 27 Fields medalist co-authors, mathandai.org Camillo De Lellis — listed in mathandai.org endorsers, checked 21 September 2026 Ravi Vakil — listed in mathandai.org endorsers, checked 21 September 2026 Timothy Gowers — declined to sign, blog post 17 September 2026 François Charles — no public position found Nikhil Srivastava — no public position found Ulrike Tillmann — no public position found Edward Witten — no public position found Melanie Matchett Wood — no public position found "Signed it" means named among the declaration's 27 Fields medalist co-authors; "endorsed it" means found in the public endorsers list. Absence is not a position.
Four of the nine have a recorded position, and they do not agree. The remaining five are absent from both lists, which is not the same as opposing the letter.

Hairer is a co-author of the declaration and a member of the group advising the company the declaration is aimed at. De Lellis and Vakil are in the endorsers list, which stood at 7,788 names when I checked this afternoon. Gowers is on the board and publicly declined. So the body OpenAI is working with contains one signatory, two endorsers, and one prominent dissenter, which is a more awkward and more credible composition than a captured board would have. As far as I can find, Hairer has not written publicly about holding both positions at once. He gave a video interview on 20 September that may touch on it; I could not get a transcript, so I am not going to characterise what he said in it.

One number to keep in mind when the 27 is quoted back at you: it was 25 at launch on 11 September, and Drinfeld and Mumford were added since. The count is a running total, not a fixed slate.

The count and the list

Now the part I find harder to be generous about.

The announcement’s opening sentence is the reason the group exists: “On August 28, we began training a new internal model. In addition to resolving the Navier–Stokes Millennium Prize problem, this model has now resolved more than 100 long-standing open problems across most areas of mathematics.”

More than 100. The next sentence links the phrase “the pace of its progress” to openai.com/index/navier-stokes-solution/#astra-internal-model-open-math-problems, and what sits at the other end is the only public evidence for the count.

The anchor lands on a chart titled “Performance of GPT-6 Astra and our Internal Model on a curated set of open math problems.” It plots two series, GPT-6 Astra and the internal model, with pass rate on the vertical axis running from 0 to 0.5 and test-time compute on the horizontal, log scale. The compute axis carries no numbers at all: ticks, but no units, no values, no range.

So the evidence offered for a hundred resolved problems is a pass rate on a curated set. Not the set. There is no list of which problems were curated, which were passed, or what counts as passing one. And the set is a benchmark, which is the specific thing 27 Fields medalists wrote a letter about ten days earlier. OpenAI’s own summary of their complaint, on the same page as the chart link, is that they “raise concerns about the negative externalities of solving open problems as a benchmark for new AI systems.”

What the Navier–Stokes page documents in full, and documents well, is one result and one side result. The Navier–Stokes resolution took roughly 88 hours of agent time from launch on 1 September to a proof on 5 September, with Lean formalization and verification adding about 17 hours. The Euler regularity disproof took “nearly 100 agents” about 50 hours. Across all attempted problems the agents sent 4.9 million messages and used about 300 billion output tokens. OpenAI says it does not intend to claim the Millennium Prize.

So of “more than 100,” two are publicly identified.

Two of at least a hundred claimed results are publicly identified A grid of one hundred squares, one per claimed resolved open problem, drawn as a floor for OpenAI's claim of "more than 100". Two squares are filled solid: the Navier–Stokes finite-time blowup result and the Euler regularity disproof, both described on OpenAI's Navier–Stokes page and both carrying a published Lean certificate in the openai/NavierStokesAndEuler repository. The remaining ninety-eight or more are drawn empty: they are not named, described, or itemised in any OpenAI page I could find. Checked 21 September 2026. CLAIMED RESULTS VS PUBLICLY IDENTIFIED RESULTS · CHECKED 21 SEP 2026 Navier–Stokes finite-time blowup — writeup, Lean certificate, Comparator challenge Euler regularity disproof — writeup, Lean certificate, Comparator challenge 2 named, written up, Lean-certified 98 or more neither named, described, nor listed anywhere public One square per claimed resolved open problem. OpenAI's wording is "more than 100", so 100 squares is a floor, not the figure — the true denominator is unpublished, which is the point. The two filled squares are Navier–Stokes and Euler. The openai/NavierStokesAndEuler repository holds exactly two Lean certificates and has two commits, the last on 10 September 2026 — eleven days before the "more than 100" claim was published.
The count is on one page. The evidence on the other is a pass rate over a curated set plus two documented results. Which problems make up the rest has not been published.

I want to be precise about what is and is not wrong here. OpenAI has not claimed that the other 98 are verified. It has claimed they are resolved, and it is asking an advisory group to help work out how to release them. Asking for advice before publishing is the correct order of operations, and is roughly what the declaration’s second complaint asked for.

But the number is already in circulation, and a number that cannot be audited is doing work that the evidence behind it cannot support. The openai/NavierStokesAndEuler repository was created on 8 September and last pushed on 10 September. I should be careful here, because the shape of it flatters my argument in a way the contents do not: the two root files, NavierStokes.lean and Euler.lean, are top-level statements sitting on top of two large proof directories, 1,839 files under Euler/ and 816 under NavierStokes/, with 2,669 in the repository as a whole. The repo’s formalization.yaml records a sorry count of zero for every main result, which means nothing is left unproved on Lean’s standard axioms. This is a serious formalization, not a gesture.

It covers two problems.

GitHub file listing for the openai/NavierStokesAndEuler repository, showing 1 Branch, 0 Tags and 2 Commits, with the directories ComparatorChallenges, Euler and NavierStokes, and the files .gitignore, Euler.lean, LICENSE and NavierStokes.lean, each last changed two weeks ago.
The top level of the verification record: two statement files, a Comparator challenge directory, and two commits. The proof directories behind them hold 2,655 files between them, and cover two of the claimed hundred-plus results. Screenshot: github.com/openai/NavierStokesAndEuler, repository contents under Apache-2.0; page chrome reproduced for commentary. Captured 21 September 2026.

The standard exists. It was applied twice.

One detail in that repository makes the gap concrete, and it is to OpenAI’s credit. Alongside the two proofs the repo ships a ComparatorChallenges directory, with instructions for checking the formalizations using Comparator, and a credit: “Thank you to the Formal Conjectures authors for their Lean formalization of the Navier–Stokes problem statement, which we adapted for these Comparator challenges.” Formal Conjectures is a Google DeepMind project. OpenAI checked its own proofs against a rival lab’s independently written statement of the problem.

That is the hard part, and I have written before about why it is the hard part: a proof checker verifies a proof against a statement and has nothing whatsoever to say about whether the statement means what you wanted. Anthropic hit the same wall formalizing Fermat’s Last Theorem and built a comparator for the same reason. Its formalization was announced on 4 September and OpenAI’s repository appeared on the 8th: two labs, four days apart, independently concluding that the statement is where the risk lives.

So OpenAI knows what an auditable mathematical claim looks like. It built one, twice, including the part nobody would have noticed if it had been skipped. The question the advisory group is going to have to answer in public is what standard applies to results three through a hundred, and whether a result with no writeup, no formalization and no name is a result or a log entry.

What has been measured, and what has not

The reason I am not sure the carve-out changes much in practice is that the pace question has almost no evidentiary basis on either side.

There is one measurement of adoption, published seventeen days before the declaration. Jin, Ke and Sui collected all 32,944 arXiv submissions with a mathematics category between 1 March and 20 August 2026 and found 1,712 that disclosed a substantive mathematical contribution from AI, a share that grew from 1.39% of mathematics submissions in March to 14.09% by 20 August. OpenAI’s systems appear most often, Anthropic’s second. The finding that belongs in this essay is their fourth: of 717 named open-problem records associated with substantive AI use, 71% are labelled fully resolved, and that label comes from the authors’ own descriptions of their work. A field-scale count of resolved open problems, resting on self-report.

What nobody has measured is the mechanism the declaration is actually worried about. That study tracks disclosure, not consequences: not how refereeing adapts, not which problems get abandoned, not whether the human transmission chain is thinning. The declaration’s central premise, that mass production at speed destroys fertile ground, remains a claim about something nobody has instrumented. It may well be true. Twenty-seven Fields medalists believing it is evidence of a kind, and not the kind that settles an argument about rates.

Nor is there a published figure for what fraction of claimed AI mathematical results have been machine-checked. Today’s example is two out of at least a hundred, but that is one company on one day, not a measured rate.

And there is no validity study for the benchmark at the centre of the whole dispute. OpenAI’s own summary of the letter says mathematicians “raise concerns about the negative externalities of solving open problems as a benchmark for new AI systems.” Nobody has tested whether solving famous open problems tracks mathematical capability, or tracks anything except itself. The field adopted the proxy and skipped the validation step, which is a pattern I keep running into from completely different directions.

So the group has been excluded from advising on a rate that nobody can currently characterise, and included on a communication problem that is tractable this week. Read uncharitably, that is a company keeping the interesting lever. Read charitably, it is the only division of labour that could produce anything useful in the near term. I lean about 60/40 toward the uncharitable reading, mostly because the carve-out did not have to be written down at all. Nobody was going to hand an outside board control of a training schedule, and saying so explicitly converts an obvious practical limit into a stated boundary.

The template both sides reached for

There is a precedent here that neither side invented, and both named.

The Leiden Manifesto for Research Metrics was published in Nature on 22 April 2015: ten principles from Diana Hicks, Paul Wouters, Ludo Waltman, Sarah de Rijcke and Ismael Rafols, written after the 2014 Leiden science-indicators conference, objecting to the use of citation counts, impact factors and h-indices as a substitute for judging research. A scientific community formalising an objection to a measurement regime that had quietly become the goal.

Tao writes that the declaration grew out of a week of discussion and that “it is unfortunate that we did not have the time to have a more consultative process, as with Leiden.” Gowers, declining to sign, writes that “it seemed better to do what I did with the Leiden Declaration and set out my own position in a blog post.” The signer and the dissenter reached for the same eleven-year-old document as their model for how this is done. That tells you they both understand the fight as being about a metric colonising a practice. It is the same fight, with the h-index replaced by the open problem.

Leiden is also a caution. Eleven years on, impact factors are still in hiring committees. Manifestos are a way of putting a position on the record, not a way of changing an incentive.

What would make the group’s first document count

The advisory group has an open input form and a stated task. Its first published recommendation is the thing to watch, and three specifics would tell us whether it has teeth:

Whether it names a verification standard for the unreleased results (writeup, formalization, external statement check, or none) and whether OpenAI accepts it. Whether it publishes the list, or asks for it. A count is not a result set, and the group cannot advise on releasing a hundred things it has not seen. And whether it says anything about pace anyway. Nothing in the charter stops it. The group may “offer advice we have not requested” and “make its advice public”; the carve-out removes OpenAI’s obligation to be advised, not the group’s freedom to advise. If the members think the rate is the problem, the charter they have already been given lets them say so on their own website, and OpenAI has committed in writing not to mind.

That is the test. Not whether the carve-out exists, but whether nine mathematicians who are not being paid decide to write past it.

References

  1. OpenAI (2026). Advisory Group on Mathematics and Artificial Intelligence. Published 21 September 2026.
  2. Advisory Group on Mathematics and Artificial Intelligence (2026). agmai.org. Hosted by the Institute for Advanced Study. Checked 21 September 2026.
  3. Avila, Bhargava, Birkar, Deligne, Deng, Donaldson, Drinfeld, Duminil-Copin, Figalli, Hairer, Huh, Kontsevich, Lindenstrauss, Lions, Maynard, McMullen, Mori, Mumford, Ngô, Okounkov, Scholze, Smirnov, Tao, Viazovska, Villani, Werner and Zelmanov (2026). A Severe Misalignment of AI in Mathematics. Published 11 September 2026; DOI 10.5281/zenodo.22737750 as stated on the page. Co-author count 27 and endorser count 7,788 as displayed on 21 September 2026.
  4. OpenAI (2026). On the Navier–Stokes Millennium Prize Problem. Describes work completed 5–8 September 2026. The linked writeup, Finite time blowup for Navier–Stokes, is a separate PDF.
  5. OpenAI (2026). openai/NavierStokesAndEuler. Lean certificates; repository created 8 September 2026, last pushed 10 September 2026.
  6. Gowers, T. (2026). Why I didn’t sign the Fields medallists’ letter. Gowers’s Weblog, 17 September 2026.
  7. Tao, T. (2026). A Severe Misalignment of AI in Mathematics. What’s new, 11 September 2026.
  8. Pachter, L. (2026). Align AI and Mathematics—to Something Else. Bits of DNA, 12 September 2026.
  9. Jin, J., Ke, Z. T. and Sui, B. (2026). The Gold Rush in AI4Math: Where Are We Now?. arXiv:2608.24961, 25 August 2026.
  10. Hicks, D., Wouters, P., Waltman, L., de Rijcke, S. and Rafols, I. (2015). Bibliometrics: The Leiden Manifesto for research metrics. Nature 520, 429–431, 22 April 2015.
  11. Google DeepMind (2026). Formal Conjectures: Lean formalization of the Navier–Stokes problem statement. Adapted by OpenAI for its Comparator challenges.
  12. Crawford, T. (2026). Navier-Stokes, AI and the Future of Mathematics with Martin Hairer (2014 Fields Medal). Tom Rocks Maths, 20 September 2026. Video; not transcribed here.