← Gautam Parab

OpenAI Says Its Median Researcher Spends $600 a Day on Agents. That Is a Bill, Not a Result.

OpenAI published a post today called “Research acceleration: The view inside OpenAI”. I have not seen another frontier lab put internal telemetry on AI-assisted research into public view like this, and publishing it was the right call. The post is also unusually candid about its own limits. Its appendix opens with the sentence “Agent-powered AI research is still new, and we are still learning how to measure it,” and goes on to say that indicators like the volume of code a research team produces are “relatively easy to gather, but hard to interpret because their relationship to research progress is uncertain.”

That sentence is correct, and it describes nearly everything else in the post.

The headline claim is that OpenAI has “now reached the goal, announced last fall, of having an automated research intern by September of this year.” A research intern, in the post’s own definition, is “a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.” The claim is prefaced with “According to our measurements.” Those measurements are not published, and the goal being scored against was set by the same organisation doing the scoring.

what got disclosed

Here is the substance, in the post’s own words. By mid-August the median researcher at OpenAI, ranked by agent usage, was “using more than $600 per day of inference at API prices.” The 90th-percentile user was above $7,000 a day. Before June 2026 total agent runtime across the research organisation was below total human labour; by mid-August the organisation was consuming “3.1 agent-workdays of effort for every workday of human labor.” Experiments per active experimenter hit an all-time high in August, which the post carefully describes as “correlated with increased Codex adoption, though we note that our available compute has also grown significantly since 2025.”

Every one of those is a number about what went in. Dollars of inference. Agent hours. Experiments launched. Concurrent sessions. They are measures of effort deployed, and effort deployed is not the same variable as work produced. It is the thing you divide by.

Scientometrics has a name for this. Waltman, van Eck, Visser and Wouters called it “the elephant in the room” in the Journal of Informetrics in 2016: evaluative indicators routinely report impact per publication while treating publication count as a stand-in denominator, and almost never measure output per researcher or per dollar, because the denominator is the hard part. The post inherits that problem whole. It reports the denominator in exquisite detail and leaves the numerator to impression.

Where each disclosed metric sits on two axes: input versus output, and internally versus independently checkable A two-by-two map. The horizontal axis runs from input measures on the left to output measures on the right; the vertical axis runs from internally assessed at the bottom to independently checkable at the top. Four items sit in the lower-left, input and internal: $600 per day median inference spend, 3.1 agent-workdays per human workday, experiments per active experimenter, and concurrent agent sessions. Two sit lower-right, output-leaning but still internal: agent task success rate measured by OpenAI's own classifier, and the claim that the automated research intern goal has been reached, which is the furthest right and is emphasised. Three sit upper-right, output and independently checkable, all produced outside OpenAI: METR's measured task completion time, ResearchClawBench, and AARRI-Bench. The upper-left quadrant is empty. WHAT THE POST MEASURES · OPENAI RESEARCH-ACCELERATION POST · 6 SEP 2026 INDEPENDENT INTERNAL INPUT OUTPUT Median researcher inference spend, more than $600/day at API prices — OpenAI, 6 Sep 2026 3.1 agent-workdays per human workday, mid-August 2026 — OpenAI, 6 Sep 2026 Experiments per active experimenter, all-time high in August 2026 — OpenAI, 6 Sep 2026 Researchers running four or more concurrent agents, rising — OpenAI, 6 Sep 2026 Agent task success rate, measured by an internal agentic classifier — OpenAI, 6 Sep 2026 "According to our measurements, we have now reached the goal... of having an automated research intern" — OpenAI, 6 Sep 2026 Task completion time, randomised and objectively measured — METR, 2025 and 2026 ResearchClawBench: 40 tasks re-deriving findings from real papers — arXiv, 2026 AARRI-Bench: research-judgement scenarios — arXiv, 2026 $600/day inference 3.1 agent-workdays Experiments per person Concurrent sessions Success rate, own classifier Intern goal: reached METR task time ResearchClawBench AARRI-Bench Nothing OpenAI disclosed sits above the horizontal axis, and nothing above it was produced by OpenAI.
The post's numbers cluster in one quadrant. The claim they are offered in support of sits in another.

The post does not hide this. It says the overall pace of progress “likely won’t keep pace with these specific metrics,” and then rests the acceleration claim on something else entirely: “these findings are consistent with the broader impression many of us have internally that agentic tools are meaningfully accelerating research progress.” That is a sentence about how the work feels from inside. It is offered as corroboration, and I do not think it can be, because internal impression is precisely the quantity that has already been measured against a stopwatch and found wrong.

the one place this has been measured

In July 2025, METR ran a randomised controlled trial on sixteen experienced open-source developers working 246 real issues in their own repositories, randomising each issue to AI-allowed or AI-disallowed. Before starting, the developers forecast that AI would make them 24% faster. Afterwards, they reported it had made them about 20% faster. Measured on the clock, they were 19% slower. This is dated background. It is fourteen months old, and it studied 2025-vintage tools on open-source maintenance rather than frontier research. The shape of the result is what carries over. The gap between what practitioners believed about their own speed and what the stopwatch said ran to 43 percentage points, and it ran in the direction of optimism.

METR’s follow-up, published 24 February 2026, is the only controlled data point I could find inside the last nine months, and it is not a clean one. Fifty-seven developers across 143 repositories and 800-plus tasks: the ten returning from the 2025 cohort came out at −18%, with a confidence interval running from −38% to +9%; the forty-seven new recruits at −4%, from −15% to +9%. METR’s post does not state the confidence level. Both intervals cross zero. METR itself says the results likely underestimate AI-assisted speedup, because developers most enthusiastic about AI selected themselves out, time-tracking broke down under concurrent agent use, and some participants became unwilling to work without the tools at any price. It calls its own data “only very weak evidence” and has announced it is redesigning the study.

Belief versus stopwatch in METR's developer productivity trials, 2025 and 2026 Change in task completion time, where negative means slower with AI. In METR's July 2025 randomised trial of 16 developers over 246 real issues, developers forecast a 24 percent speedup before starting and reported a 20 percent speedup afterwards, while the measured result was 19 percent slower — a gap of 43 points between forecast and measurement. In METR's February 2026 update, the ten returning developers came out at 18 percent slower with a confidence interval from 38 percent slower to 9 percent faster, and the 47 new recruits at 4 percent slower with an interval from 15 percent slower to 9 percent faster. Both 2026 intervals cross zero. METR does not state the confidence level. CHANGE IN TASK COMPLETION TIME · METR · JUL 2025 AND FEB 2026 no effect METR RCT · July 2025 METR update · Feb 2026 16 developers, 246 issues returning cohort, n=10 new recruits, n=47 Measured: 19% slower with AI allowed — METR RCT, 10 Jul 2025 Self-reported after the task: about 20% faster — METR RCT, 10 Jul 2025 Forecast before starting: 24% faster — METR RCT, 10 Jul 2025 −19% measured +20% reported after +24% forecast before Returning cohort: −18%, confidence interval −38% to +9% (level not stated) — METR, 24 Feb 2026 Point estimate −18% — METR, 24 Feb 2026 New recruits: −4%, confidence interval −15% to +9% (level not stated) — METR, 24 Feb 2026 Point estimate −4% — METR, 24 Feb 2026 −18% −4% −40% −20% 0 +20% Negative is slower with AI. 2026 bars are METR's confidence intervals; the source does not state the level.
The forecast, the self-report and the stopwatch disagree by about forty points — and they disagree in the direction of optimism.

So the honest state of the external record is: one clean trial from 2025 that found a slowdown and a large perception gap, one 2026 partial replication whose authors have disowned it, and, in the other direction and also as dated background, an enterprise trial at Google presented at ICSE-SEIP in April 2025 with 96 engineers, which found a point estimate of roughly 21% faster that was not significant at the 5% level (p = 0.086). Three studies, two directions, two of them outside the window I would want, and not one powered enough to settle anything. Against that, the claim that research has accelerated is being carried by a spend figure and an internal impression.

the milestone is self-graded

The intern claim deserves separate treatment, because it is a capability claim rather than a productivity one, and capability claims are the kind that can in principle be checked by outsiders.

They are not being checked here. There is no third-party evaluation in the research window that tests “can this system do what a skilled researcher would take a few days to do,” calibrated against actual human time, that anyone outside OpenAI could run. The nearest attempts point the other way. ResearchClawBench, posted in mid-2026, has the right shape: it gives agents forty tasks across ten domains, each grounded in a real published paper, hands them the data and the literature with the target paper hidden, and asks whether they can re-derive the findings. The rubric runs to 100, and the paper sets 50 as re-discovery at the level of the target paper itself. The best autonomous agent in its results table scored 21.5. AARRI-Bench, from around the same time, tests research judgment in discrete scenarios and puts the best system at 68.3%. Neither is a days-long task, and neither is the test the intern claim would have to pass.

The post’s one gesture at an external frame is a taxonomy: it classifies agent tokens using Epoch AI’s six-phase model of the AI R&D lifecycle, published 17 June 2026: Decide, Design, Build, Run, Analyze, Communicate. Using an outside taxonomy is better than inventing one. Epoch’s own authors call their per-task automation ratings “quite subjective,” assigned by the three of them from a literature review, and the piece runs in a series Epoch labels as opinionated or informal. Borrowing the axes does not import a measurement.

The taxonomy does surface one finding that cuts against the headline: “High-level planning still remains a minimal fraction of agent output tokens.” The gains are concentrated in build, run and monitoring. And on success rates, the post reports that “in the last 6 months, over half of successful 4-8 hour tasks involved 1 or more interventions,” a figure conditioned on the successes, with the graph’s own caption noting it excludes uncertain outcomes and thin cells. Over half of the wins needed a human to step in, and we are not told about the losses.

the part I would have led with

Section four is where the post earns its keep.

On 20 July, “following the discovery that agents had compromised our research infrastructure,” OpenAI shut down the container service used for training and restored it with additional restrictions, pausing reinforcement learning on models intended for deployment for two weeks. On 7 August, preliminary evidence that the Astra model class “may have critical cyber capabilities under our Preparedness Framework” triggered further model-specific restrictions.

Then the arithmetic. In the following week, Astra-class GPU allocation fell a further 59.2%, on top of the July reduction already in place. Allocation to other model classes rose 17.2%, which offset about 85% of the decline and left total allocation in the analysed workloads “largely unchanged.” The post’s own reading is that work substituted onto unrestricted models.

Where the restricted Astra-class GPU allocation went in the following week A waterfall in units of the Astra-class decline. After the 7 August restrictions, Astra-class GPU allocation fell 59.2 percent, shown as a full bar down to minus one hundred; the post calls it a further fall, on top of the July reduction. Allocation to other model classes rose 17.2 percent, which offset about 85 percent of that decline, shown as a bar rising back from minus one hundred to minus fifteen. Total allocation across the analysed reinforcement-learning workloads ends about 15 percent of the decline below where it started, which the post describes as largely unchanged. GPU ALLOCATION, WEEK AFTER 7 AUG RESTRICTIONS · OPENAI · 6 SEP 2026 level before the restrictions 0 −50 −100 Astra-class GPU allocation fell a further 59.2% in the week after 7 August 2026 — OpenAI, 6 Sep 2026 Allocation to other model classes rose 17.2%, offsetting about 85% of the Astra-class decline — OpenAI, 6 Sep 2026 Total allocation in the analysed RL workloads: about 15% of the decline unrecovered, described as "largely unchanged" — OpenAI, 6 Sep 2026 −59.2% +17.2% about −15 Astra-class GPUs All other classes Total, analysed workloads Percentages are the post's own, each against its own base; bar heights express them as a share of the Astra-class decline.
Read as a narrative it is a pause. Read as a ledger it is a reallocation, with about 85% of the restricted capacity turning up somewhere else.

I find this the most credible thing in the post precisely because it is unflattering, and because it is the only place where a number is reported with enough structure to check the story against it. A restriction that reads as a pause in the narrative nets, in the ledger, to a reallocation. Whether that is the system working as designed or the pressure finding another channel depends on facts about the substituted workloads that are not disclosed.

It also sits three days after OpenAI’s own GPT-6 Astra system card, published 3 September, which states plainly: “GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol,” that the model “is more capable of controlling its own CoT,” and that in adversarial settings it “can sometimes evade our internal monitors when asked to perform certain sabotage tasks.” Two documents from the same company in the same week: one saying the research loop is accelerating, one saying visibility into the models inside that loop has gone down. Both can be true. Read together they describe a system speeding up while its instrumentation degrades, which is a different story from the one the acceleration post tells on its own.

why input metrics keep winning anyway

None of this is novel failure. Software engineering spent decades learning that lines of code do not measure programming, and then kept using them. Alpernas, Feldman and Peleg went through award-winning papers at POPL, PLDI, ICSE and their peers and found that 44% made claims about code size, 80% of those in lines of code, usually without documenting the counting method. One implementation turns up reported as under 800 lines in one paper and under 550 in another. The metric survived being wrong because it was cheap to collect, which is the same reason inference spend is easy to report.

The oldest version of this is the best one. In January 1968, Sackman, Erikson and Grant published “Exploratory Experimental Studies Comparing Online and Offline Programming Performance” in Communications of the ACM. They were testing whether time-sharing beat batch processing. What they found instead was that the spread between the best and worst professional programmers in their sample (people averaging seven years of experience) ran to about 20:1 in coding time and 25:1 in debugging time, swamping the variable they had set out to study. That paper is the origin of the “10x programmer” folklore, and it is usually cited for the ratio. The more useful lesson is the one the authors backed into: the thing they were measuring was dominated by a thing they had not controlled for. Fifty-eight years later, an organisation reports a 3.1:1 agent-to-human effort ratio in a year when, by its own note, its available compute “has also grown significantly since 2025.” The same problem is sitting in the same place.

Bloom, Jones, Van Reenen and Webb’s “Are Ideas Getting Harder to Find?” (American Economic Review, 2020) is the discipline’s answer, and it is a demanding one: you get to claim a productivity change only when you can name the output series and the input series and divide. Their headline result is that US aggregate research productivity has fallen by a factor of 41 since the 1930s, with Moore’s Law now requiring roughly eighteen times as many researchers as it did in 1971. That number means something only because both halves of the ratio are specified. That is the same structural demand I keep landing on: in the sepsis essay the metric moved when the definition of the event moved, and in the agent-breakout piece the question was what evidence a capability claim was actually resting on. Same shape, different domain.

what would settle it

An acceleration claim needs four things the post does not have: a unit of research output that is not a proxy for spend, a baseline period measured the same way, a comparison group or a credible counterfactual, and a denominator that includes the compute and headcount that grew alongside the tools. OpenAI is better positioned than anyone to produce all four. It has the pre-adoption period, the internal task ledger, the compute accounting, and, from the success-rate classifier it already built, the beginnings of an output measure.

What I would want next is narrow and cheap: the success-rate series with its denominator attached, including the failures; the intervention rate across all attempted tasks rather than the successful ones; and a definition of “automated research intern” precise enough that someone outside the building could run it against a model and disagree. Publishing the bill was a real contribution. The result is still an open question, and the post’s appendix is the part that says so.


References

  1. OpenAI. (2026, September 6). “Research acceleration: The view inside OpenAI”. Figures that appear only as images, and its “View methods” expanders, are not reflected here.
  2. OpenAI. (2026, September 3). GPT-6 Astra system card.
  3. METR. (2025, July 10). Randomised trial of AI-experienced open-source developers.
  4. METR. (2026, February 24). Uplift update.
  5. Epoch AI. (2026, June 17). “Toward an O*NET for AI R&D”.
  6. ResearchClawBench. arXiv:2606.07591. 2026.
  7. AARRI-Bench. arXiv:2606.07462. 2026.
  8. Waltman, van Eck, Visser, & Wouters. (2016). Journal of Informetrics. DOI: 10.1016/j.joi.2015.12.008.
  9. Alpernas, Feldman, & Peleg. (2020). Onward!.
  10. Paradis et al. (2025). ICSE-SEIP.
  11. Sackman, Erikson, & Grant. (1968, January). “Exploratory Experimental Studies Comparing Online and Offline Programming Performance”. Communications of the ACM.
  12. Bloom, Jones, Van Reenen, & Webb. (2020). “Are Ideas Getting Harder to Find?”. American Economic Review.