In July 1966, MIT AI Memo AIM-100, written by Seymour Papert, proposed the Summer Vision Project: “an attempt to use our summer workers effectively in the construction of a significant part of a visual system.” The plan was to have a group of largely undergraduate summer students solve figure-ground segmentation, dividing a camera image into likely objects, likely background, and chaos, over one summer. It is a six-page document, still online. Computer vision took roughly four decades longer than the memo budgeted.
I keep returning to it because of how specific it is. Not a vague prophecy about thinking machines but a staffing plan with a deadline, made by people with full access to the state of the art, about a problem they had already decomposed. It is the most precisely dated capability misestimate in the field.
That is the discount I apply to the forecasts now in circulation. The most-cited expert survey remains Grace et al., “Thousands of AI Authors on the Future of AI”: 2,778 researchers who had published at top AI venues put the chance of unaided machines outperforming humans at every possible task at 10% by 2027 and 50% by 2047. Prediction-market-style aggregates run earlier; Metaculus’s long-running question on the first general AI system has sat in the early 2030s recently, though that figure moves and is worth checking live rather than quoting from an essay. The memo does not establish that timelines are inherently too optimistic. It establishes that the people best positioned to estimate were confidently wrong about something they had already broken into parts, and I hold current forecasts to the same discount.
What survives the discount is the instruments. There is no single measurement of progress toward AGI, and anyone who quotes you one number is quoting an instrument, not a phenomenon. What exists is three families: benchmarks that measure difficulty (can it answer harder questions?), benchmarks that measure duration (how long a task can it finish unsupervised?), and benchmarks that measure economic work (would a professional accept the deliverable?). Each has a different failure mode. Two of them have documented answer-key error rates high enough to swallow a year of apparent progress. And the axis that determines whether any of it is usable, cost per task, is almost never in the headline.
Difficulty benchmarks have a structural problem: they are designed to be beaten, and once beaten they tell you nothing. Duration is different, because the quantity being measured, how long a task a system can carry alone, is the same quantity a manager thinks in.
METR’s time horizon work operationalizes this: the 50%-task-completion time horizon is the length of task, measured by how long humans take on it, that a model completes autonomously with 50% success. The original paper (Kwa, West, Becker et al., March 2025) reported that this horizon had doubled roughly every seven months since 2019.
The number that matters more is in the Time Horizon 1.1 update from 29 January 2026, which expanded the suite from 170 to 228 tasks. The doubling time is not constant. On the all-time trend it is 196.5 days. Measured since 2023 it is 130.8 days. Measured since 2024 it is 88.6 days. The trend line people cite as evidence of steady progress is, on METR’s own data, bending.
METR is unusually honest about what this does not mean. Its limitations note states that error bars are about a factor of two in each direction, that horizons differ between domains by orders of magnitude, with visual computer-use tasks landing 40 to 100 times lower than software tasks, and, most importantly, that a 50% time horizon of X hours does not mean you can delegate X-hour tasks. Half the time it fails, and most real work needs far better than a coin flip. The time-horizons page adds that measurements above 16 hours are unreliable with the current task suite. That is a ceiling on the instrument, not on the models.
Two of the most-cited difficulty benchmarks have published, quantified problems with their own answer keys, which changed how I read benchmark deltas.
Epoch AI’s FrontierMath, 338 problems, 295 in Tiers 1–3 and 43 research-level problems in Tier 4, released a v2 on 12 June 2026 “addressing errors in 42% of problems,” correcting 123 in Tiers 1–3 and 12 in Tier 4. That is the benchmark maintainers’ own disclosure, and it is to their credit that they published it. But it means every cross-version score comparison on that benchmark is comparing two different tests.
Humanity’s Last Exam, the 2,500-question CAIS and Scale AI benchmark introduced in January 2025, has a similar problem, found externally. FutureHouse’s audit of 321 text-only chemistry and biology questions concluded that “29 ± 3.7% (95% CI) of the text-only chemistry and biology questions had answers with directly conflicting evidence in peer reviewed literature.”
I want to be careful here: neither of these makes the benchmarks useless. A test with a 29% key-error rate still separates a model that scores 5% from one that scores 45%. What it cannot do is adjudicate a three-point gap between two frontier models, which is precisely what it is most often used for. This is the same failure mode I described when looking at what the text-to-SQL benchmarks actually show: the benchmark measures the benchmark, and the leap to “can it do the job” is made by the reader, not the data.
ARC-AGI is the one major benchmark that treats spend as a first-class measurement rather than a footnote. Its 2025 results analysis reports verified ARC-AGI-2 scores alongside dollars per task, and the picture that emerges is a purchase decision, not a ladder of intelligence.
The human baseline matters too, and ARC’s is concrete: in a calibration study with over 400 participants, every ARC-AGI-2 task was solved by at least two humans within two attempts. So the benchmark’s ceiling is a thing people already do, not an aspirational target.
The last failure mode is the leaderboard rather than the test. “The Leaderboard Illusion” (Singh, Nan, Kapoor, Longpre, Hooker et al., April 2025) audited Chatbot Arena and found that a small number of large providers received roughly 19–20% each of total battle data while 83 open-weight models shared about 29.7% between them, that private variants could be tested and selectively disclosed, and that access to Arena data yielded Arena-specific gains of up to 112%. No one cheated. A leaderboard with unequal sampling and selective disclosure measures participation as much as quality, the same lesson, in a different domain, as the consultancy surveys on AI profitability, where the instrument was the phenomenon.
Against difficulty and leaderboards, OpenAI’s GDPval (October 2025) attempts a different question: real deliverables across 44 occupations, built from work by professionals averaging 14 years of experience, graded by expert preference. Its reported win-plus-tie rates against human experts at publication, Claude Opus 4.1 around 47.6% and GPT-5 around 39%, are the only widely cited numbers I know of that answer “would a professional take this?” They are also, by now, nearly a year stale, and I have not found a refreshed 2026 run.
Four things survive all of that, in this order. Time horizon, but only the order of magnitude, never the ranking between this year’s models, because their intervals overlap. Cost per task alongside every accuracy number, because unpriced state-of-the-art is a marketing claim. Refreshed economic-work evaluations, because those are the only ones whose units translate to a decision I might make. And the benchmark changelogs, because a test that corrected 42% of its own problems will do it again.
What does not survive is any single number offered as progress toward AGI. There isn’t one. There is a set of instruments with known biases, and the honest version of the question is not “how close are we” but “which instrument, measuring what, at what cost, with what error bar.”
References
Where a number is a lab’s own report of its own system, it is labeled as such. Aggregator-sourced 2026 leaderboard figures circulating for several benchmarks could not be traced to a primary source and were left out.