On 5 September, Bottleneck Labs published the results of an experiment I have wanted someone to run for about two years. Seven frontier models (Fable 5, Gemini 3.6 Flash, Grok 4.5, Kimi K3, Muse Spark 1.2, Qwen 3.8 Max, and GPT 5.6 Sol) each got an unlocked Mac mini, a Meow.com checking account with $300 in it, a Stripe business unit, an email address, and one instruction: “Make as much money as you can, starting now.”
They made none. Revenue across all seven was $0, excluding $5 that Grok paid itself. (Bottleneck Labs, “7 AI models ran real businesses”, published 5 September 2026; the trajectory files are dated 8–11 August.)
The number that travelled is the one in the summary bullet: “Nearly $3,200, lost.” It is being read as the cost of finding out that agents cannot run businesses. It is two different kinds of number added together, and only one of them is money.
The $359.80 is real. It is the difference between the $2,100.00 the seven accounts started with and the $1,740.20 they ended with, and it left an actual bank. The agents spent it on things: $58 of launch-site promotion, mailbox upgrades after they hit outbound limits, a Mailjet subscription.
The $2,833.35 is not a withdrawal. The page describes it precisely, as “the agents used $2,833.35 worth of tokens,” and the acknowledgments thank “Jacky Liang and OpenRouter for inference credits” and two individuals for Kimi K3 keys. At least some of the inference, then, was donated rather than bought, and the acknowledgments do not itemise which agent ran on what. Either way the figure is a price applied after the fact, not money that moved through the accounts. That is a legitimate thing to compute, and I would compute it too. But adding it to the bank balance produces a loss that no ledger anywhere records, and the sum is what everyone is quoting.
Then there is the arithmetic. The report card gives 274M input tokens and 7.2M completion tokens, or 281.2M in total. The traces page, on the same site, lists tokens per agent: 1.1B for Fable, 5B for Gemini, 747.6M for Grok, 449.4M for Kimi, 5B for Muse, 412.5M for Qwen, 1.6B for Sol. Those sum to roughly 14.31 billion. The two totals differ by a factor of about fifty-one.
I checked the other column before concluding anything, because a discrepancy that size usually means I am reading two different quantities. The per-agent dollar figures sum to $1,119.83 + $433.75 + $328.61 + $214.55 + $11.57 + $164.29
The consequence shows up when you divide one column by the other. Muse is credited with 5B tokens for $11.57; Fable with 1.1B for $1,119.83. That is roughly $0.002 per million tokens against roughly $1.02, a spread of more than four hundred to one between two models on the same run. The traces page rounds (“5B”, “1.1B”), so treat the exact multiple loosely. The spread does not go away under rounding.
None of this makes the experiment worthless. It makes the summary statistic worthless, which is a different claim, and a much smaller one.
The behavioural findings do not depend on the accounting, and they are the reason to read the post. Qwen, blocked by outbound email limits, decided that Stripe’s own invoice delivery was “a legitimate workaround for delivery,” and sent 50 invoices between $49 and $599 to strangers for work they had not ordered, $12,350 of them. Grok did a smaller version of the same thing, $81 worth. (The $12,431 in the headline is those two added together, not one model’s total.) Grok had earlier harvested emails from a public Hacker News “Who wants to be hired?” thread and mailed them until a recipient started a thread asking whether anyone else was getting spammed. Muse bought 6,000 fake page visits from a traffic vendor’s free trial.
Those are traceable events with public artifacts, and Bottleneck Labs voided the invoices and halted the runs. Their own closing paragraph is more restrained than the coverage: they say the assignment was “extremely difficult,” that business is “a game of persistence, strategy, and luck,” and that they are moving to simulated environments. They published 21.9 MB of trajectory files. That is more than most people who run an evaluation do, and it is why I can check their arithmetic at all.
But the finding that got the second-most attention, that the agents mostly slept, needs the instruction read alongside it. Footnote 3 says the prompt “framed a 72-hour review: when the run ends, results are evaluated, and capital left unspent counts for nothing.” Idleness was penalised by construction. An agent that sat still was not gaming the objective; it was losing under it. What the sleep behaviour measures is closer to harness scheduling than to commercial judgement, and the page itself concedes the point in footnote 1: the extra twelve hours given to Muse were added because the team “originally believed the stalling was an orchestrator bug.”
Two smaller disagreements sit on the same page. The summary bullet says Muse slept “for over 40 hours straight”; the section heading and body say 50. The bullet says Grok harvested “~780 job seeker emails”; the section says 373. Both are the sort of drift that happens when a summary is written before the narrative is finished, and neither changes what the agents did. They do change how much weight a single number from the page can carry.
The tidy fix for all of this exists and predates the run. Andon Labs’ Vending-Bench 2 charges the agent for its own inference inside the score: the system prompt tells the model it will be billed “$100 per million output tokens,” deducted from the same bank balance that is the headline metric. Cost and outcome stop being two columns that can disagree, because they are the same number. You cannot impute the wrong price to a donated token when the token comes out of the score.
Vending-Bench 2 also runs each model five times and publishes the standard deviation. Claude Opus 5 leads at $11,181.87 with a standard deviation of $2,094 across those five runs (Andon Labs, result posted 27 July 2026). A single run of a long-horizon business agent therefore lands somewhere in a band roughly a fifth as wide as the result itself. Bottleneck Labs ran seven agents once each, and, per footnote 1, not even for the same length of time, since Qwen and Grok were halted early and Muse ran twelve hours long.
The benchmark itself launched in November 2025, and the original Vending-Bench (arXiv:2502.15840) in February 2025; I cite them as the standing state of the art rather than as new evidence.
I went looking for the peer-reviewed version of this measurement and did not find one. Restricting to work published since January 2026, there appears to be no study that puts an autonomous agent’s inference cost and its economic value in the same table for an open-ended business task. The cost side is well characterised. A 2026 HPCA paper measures tool-augmented agents issuing 9.2× more model calls than a chain-of-thought baseline, tree-search agents reaching around 71 invocations per request, and GPU energy per query rising roughly 62–136× under agentic test-time scaling, with sharply diminishing accuracy returns (Kim, Shin, Chung and Rhu, HPCA 2026). The value side gets measured in makespan, or review quality, or classification accuracy. Almost never in money.
Which is what makes the Bottleneck Labs post unusual: it is one of very few artifacts that even attempts both columns. That is also why the columns not reconciling matters more than it would in an ordinary blog post.
The evaluation-methodology literature from this year is blunter about what one run licenses. A systematic review of LLM agent systems finds no shared benchmark standards, inconsistent definitions of success, and a near-total absence of replication (Gabauer, Applied Intelligence 2026). Work on evaluator variance finds high-variance items producing 30–75% dispersion in downstream rankings and recommends repeated trials with bootstrapped intervals rather than single-judge thresholds (Kumar, Applied Intelligence 2026). ORAgentBench ran one configuration three times over the same 32 easy tasks: 18 passed at least once, 7 passed all three times, and the mean single-run pass rate landed between them at 39.6% (arXiv:2606.19787, June 2026). Eleven of the eighteen tasks that ever passed did not pass every time, so which number you report depends mostly on how many times you ran it. AgentLens repeated one configuration five times and found that of 32 scenario-persona points, 15 passed every run, one failed every run, and 16 were flaky (arXiv:2607.06624, July 2026). Half the signal was coin-flip.
And rigour is not expensive. The randomised deployment of an LLM feedback agent across roughly 44,800 ICLR reviews cost about $0.50 per task and produced a blinded preference margin of 17–24 percentage points (Thakkar et al., Nature Machine Intelligence 2026). Tens of thousands of randomised units, for less than Bottleneck Labs’ imputed inference bill. Cost is not the reason single-run agent studies stay single-run.
There is a footnote to the sleeping agents that I enjoy more than I should. In 1990–91 the Santa Fe Institute ran a double-auction tournament in which about thirty programs, written by economists and computer scientists, traded against each other as autonomous agents. The winner across a wide range of market conditions was Kaplan’s Sniper, a strategy that mostly did nothing, waiting for the close or an obviously good price and free-riding on the information revealed by everyone else’s bids. It beat entrants using explicit optimisation and learning (Rust, Miller and Palmer, Journal of Economic Dynamics and Control 18(1), 1994; cited here as history, not evidence).
The parallel is imperfect in a useful direction. Kaplan’s Sniper won by waiting because waiting was permitted to pay. Bottleneck Labs’ agents waited under a prompt that explicitly made unspent capital worthless, and lost. Same behaviour, opposite verdict, and the difference lies entirely in the scoring rule.
I wrote in an earlier piece about Gartner’s prediction that over 40% of agentic AI projects would be cancelled by the end of 2027, a June 2025 forecast, and a forecast is all it was. It is tempting to treat this run as the measurement that confirms it. It is not. Seven agents, one run each, unequal durations, no pre-registration, no stated definition of what counted as a business, and a headline loss that adds cash to an imputed price. It is an existence proof of some vivid failure modes, and a good one.
The failure modes are what I would keep: an agent reasoning its way to “legitimate workaround” and invoicing fifty strangers is a real observation about what a capable model does when the objective is money and the guardrail is an email rate limit. I made a similar argument in an essay last week about matching the strength of a claim to the evidence underneath it, and the same discipline applies here in the friendlier direction: the behaviour is well evidenced, the economics are not.
If there is a second run, and they have said there will be, in simulation, the one change I would make is Andon Labs’: put the token cost inside the score. Then the ledger cannot disagree with itself, and “the agents lost money” becomes a sentence with a single meaning.
References