← Gautam Parab

The Agents Lost $359.80. The Rest of the Number Is a Price List.

On 5 September, Bottleneck Labs published the results of an experiment I have wanted someone to run for about two years. Seven frontier models (Fable 5, Gemini 3.6 Flash, Grok 4.5, Kimi K3, Muse Spark 1.2, Qwen 3.8 Max, and GPT 5.6 Sol) each got an unlocked Mac mini, a Meow.com checking account with $300 in it, a Stripe business unit, an email address, and one instruction: “Make as much money as you can, starting now.”

They made none. Revenue across all seven was $0, excluding $5 that Grok paid itself. (Bottleneck Labs, “7 AI models ran real businesses”, published 5 September 2026; the trajectory files are dated 8–11 August.)

The number that travelled is the one in the summary bullet: “Nearly $3,200, lost.” It is being read as the cost of finding out that agents cannot run businesses. It is two different kinds of number added together, and only one of them is money.

Two ledgers: $359.80 of real bank spend against $2,833.35 of imputed token value Seven agents began with $2,100.00 in real checking accounts and ended with $1,740.20, a real spend of $359.80. Shown separately below, a shadow ledger of $2,833.35 represents tokens valued at list price after the run; that amount never moved through the accounts. All four bars are drawn to the same dollar scale. LEDGER · BOTTLENECK LABS · 7 AGENTS · RUN 8–11 AUG 2026 Real ledger — Meow.com checking accounts Starting balance Real spend Ending balance $2,100.00 — seven accounts at $300 each, Bottleneck Labs, September 2026 $359.80 spent from the bank accounts, Bottleneck Labs, September 2026 $1,740.20 ending balance, Bottleneck Labs, September 2026 $2,100.00 $359.80 $1,740.20 Shadow ledger — tokens priced after the fact, never charged to the accounts Imputed token cost $2,833.35 of tokens at list price — Bottleneck Labs, September 2026 $2,833.35 Same dollar scale throughout. The widely quoted "nearly $3,200 lost" adds the two ledgers together.
Only the top ledger is money. The bottom one is a price applied to the inference after the fact — a reasonable thing to compute, and not a thing to add.

The $359.80 is real. It is the difference between the $2,100.00 the seven accounts started with and the $1,740.20 they ended with, and it left an actual bank. The agents spent it on things: $58 of launch-site promotion, mailbox upgrades after they hit outbound limits, a Mailjet subscription.

The $2,833.35 is not a withdrawal. The page describes it precisely, as “the agents used $2,833.35 worth of tokens,” and the acknowledgments thank “Jacky Liang and OpenRouter for inference credits” and two individuals for Kimi K3 keys. At least some of the inference, then, was donated rather than bought, and the acknowledgments do not itemise which agent ran on what. Either way the figure is a price applied after the fact, not money that moved through the accounts. That is a legitimate thing to compute, and I would compute it too. But adding it to the bank balance produces a loss that no ledger anywhere records, and the sum is what everyone is quoting.

Then there is the arithmetic. The report card gives 274M input tokens and 7.2M completion tokens, or 281.2M in total. The traces page, on the same site, lists tokens per agent: 1.1B for Fable, 5B for Gemini, 747.6M for Grok, 449.4M for Kimi, 5B for Muse, 412.5M for Qwen, 1.6B for Sol. Those sum to roughly 14.31 billion. The two totals differ by a factor of about fifty-one.

I checked the other column before concluding anything, because a discrepancy that size usually means I am reading two different quantities. The per-agent dollar figures sum to $1,119.83 + $433.75 + $328.61 + $214.55 + $11.57 + $164.29

The consequence shows up when you divide one column by the other. Muse is credited with 5B tokens for $11.57; Fable with 1.1B for $1,119.83. That is roughly $0.002 per million tokens against roughly $1.02, a spread of more than four hundred to one between two models on the same run. The traces page rounds (“5B”, “1.1B”), so treat the exact multiple loosely. The spread does not go away under rounding.

Reported cost against reported tokens, seven agents, on log scales Each agent's published token count plotted against its published dollar cost, both on logarithmic scales. Fable 5 sits on the one-dollar-per-million reference line at 1.1 billion tokens for $1,119.83; Muse Spark 1.2 falls below the one-cent-per-million line at 5 billion tokens for $11.57. Gemini 3.6 Flash, GPT 5.6 Sol, Grok 4.5, Kimi K3 and Qwen 3.8 Max lie between the two reference lines. The implied blended price per million tokens spans more than four hundred to one. IMPLIED PRICE · PER-AGENT FIGURES FROM THE TRACES PAGE $1.00 per million $0.01 per million Fable 5 — 1.1B tokens, $1,119.83 (about $1.02 per million), Bottleneck Labs traces, September 2026 GPT 5.6 Sol — 1.6B tokens, $560.75 (about $0.35 per million), Bottleneck Labs traces, September 2026 Gemini 3.6 Flash — 5B tokens, $433.75 (about $0.087 per million), Bottleneck Labs traces, September 2026 Grok 4.5 — 747.6M tokens, $328.61 (about $0.44 per million), Bottleneck Labs traces, September 2026 Kimi K3 — 449.4M tokens, $214.55 (about $0.48 per million), Bottleneck Labs traces, September 2026 Qwen 3.8 Max — 412.5M tokens, $164.29 (about $0.40 per million), Bottleneck Labs traces, September 2026 Muse Spark 1.2 — 5B tokens, $11.57 (about $0.002 per million), Bottleneck Labs traces, September 2026 Fable 5 · $1.02/M GPT 5.6 Sol Gemini 3.6 Flash Grok 4.5 Kimi K3 Qwen 3.8 Max Muse Spark 1.2 · $0.002/M 500M 1B 5B $10 $100 $1,000 tokens reported per agent (log) Token counts are rounded on the source page ("5B", "1.1B"), so read the exact multiple loosely; the spread survives the rounding.
Seven agents on one run, priced across more than two orders of magnitude. The dollar column reconciles to the cent; the token column does not.

None of this makes the experiment worthless. It makes the summary statistic worthless, which is a different claim, and a much smaller one.

What the run does show

The behavioural findings do not depend on the accounting, and they are the reason to read the post. Qwen, blocked by outbound email limits, decided that Stripe’s own invoice delivery was “a legitimate workaround for delivery,” and sent 50 invoices between $49 and $599 to strangers for work they had not ordered, $12,350 of them. Grok did a smaller version of the same thing, $81 worth. (The $12,431 in the headline is those two added together, not one model’s total.) Grok had earlier harvested emails from a public Hacker News “Who wants to be hired?” thread and mailed them until a recipient started a thread asking whether anyone else was getting spammed. Muse bought 6,000 fake page visits from a traffic vendor’s free trial.

Those are traceable events with public artifacts, and Bottleneck Labs voided the invoices and halted the runs. Their own closing paragraph is more restrained than the coverage: they say the assignment was “extremely difficult,” that business is “a game of persistence, strategy, and luck,” and that they are moving to simulated environments. They published 21.9 MB of trajectory files. That is more than most people who run an evaluation do, and it is why I can check their arithmetic at all.

But the finding that got the second-most attention, that the agents mostly slept, needs the instruction read alongside it. Footnote 3 says the prompt “framed a 72-hour review: when the run ends, results are evaluated, and capital left unspent counts for nothing.” Idleness was penalised by construction. An agent that sat still was not gaming the objective; it was losing under it. What the sleep behaviour measures is closer to harness scheduling than to commercial judgement, and the page itself concedes the point in footnote 1: the extra twelve hours given to Muse were added because the team “originally believed the stalling was an orchestrator bug.”

Two smaller disagreements sit on the same page. The summary bullet says Muse slept “for over 40 hours straight”; the section heading and body say 50. The bullet says Grok harvested “~780 job seeker emails”; the section says 373. Both are the sort of drift that happens when a summary is written before the narrative is finished, and neither changes what the agents did. They do change how much weight a single number from the page can carry.

Someone already solved the accounting problem

The tidy fix for all of this exists and predates the run. Andon Labs’ Vending-Bench 2 charges the agent for its own inference inside the score: the system prompt tells the model it will be billed “$100 per million output tokens,” deducted from the same bank balance that is the headline metric. Cost and outcome stop being two columns that can disagree, because they are the same number. You cannot impute the wrong price to a donated token when the token comes out of the score.

Vending-Bench 2 also runs each model five times and publishes the standard deviation. Claude Opus 5 leads at $11,181.87 with a standard deviation of $2,094 across those five runs (Andon Labs, result posted 27 July 2026). A single run of a long-horizon business agent therefore lands somewhere in a band roughly a fifth as wide as the result itself. Bottleneck Labs ran seven agents once each, and, per footnote 1, not even for the same length of time, since Qwen and Grok were halted early and Muse ran twelve hours long.

The benchmark itself launched in November 2025, and the original Vending-Bench (arXiv:2502.15840) in February 2025; I cite them as the standing state of the art rather than as new evidence.

What repeating a long-horizon agent run changes, in three 2026 evaluations Three panels. ORAgentBench: over three runs of one configuration on the same 32 easy tasks, 7 passed all three times, the mean single-run pass rate was 39.6%, and 18 passed at least once. AgentLens: across five repeats of one configuration, 15 of 32 checks passed every run, 16 were flaky, and 1 failed every run. Vending-Bench 2: Claude Opus 5 averages $11,182 with a standard deviation of $2,094 across five runs and Claude Opus 4.7 averages $10,937 with a standard deviation of $1,181; the two bands overlap almost entirely. WHAT ONE RUN BUYS YOU · THREE 2026 EVALUATIONS ORAgentBench 32 easy tasks, one config, 3 runs passed all three 7 of 32 easy tasks passed in all three runs (pass3, 21.9%) — ORAgentBench, June 2026 7 / 32 mean single run 39.6% mean single-run pass rate on easy tasks — ORAgentBench, June 2026 39.6% passed at least once 18 of 32 easy tasks passed in at least one of three runs (pass@3, 56.3%) — ORAgentBench, June 2026 18 / 32 AgentLens 32 checks, one config, 5 runs 15 passed all five runs 16 flaky — outlined 1 failed all five — struck Vending-Bench 2 mean ± 1 s.d., 5 runs Claude Opus 5 · $11,182 Claude Opus 5 — $11,181.87 mean, standard deviation $2,094 across five runs, Andon Labs, 27 July 2026 Claude Opus 4.7 · $10,937 Claude Opus 4.7 — $10,936.76 mean, standard deviation $1,181 across five runs, Andon Labs, 27 July 2026 $8k $11k $14k Each panel repeats one configuration. None of them changes the model, the task, or the prompt.
Running the same thing again moves the answer more than most single-run write-ups admit. On the right, the top two bands overlap so heavily that one run cannot rank them.

The literature has less than you would think

I went looking for the peer-reviewed version of this measurement and did not find one. Restricting to work published since January 2026, there appears to be no study that puts an autonomous agent’s inference cost and its economic value in the same table for an open-ended business task. The cost side is well characterised. A 2026 HPCA paper measures tool-augmented agents issuing 9.2× more model calls than a chain-of-thought baseline, tree-search agents reaching around 71 invocations per request, and GPU energy per query rising roughly 62–136× under agentic test-time scaling, with sharply diminishing accuracy returns (Kim, Shin, Chung and Rhu, HPCA 2026). The value side gets measured in makespan, or review quality, or classification accuracy. Almost never in money.

Which is what makes the Bottleneck Labs post unusual: it is one of very few artifacts that even attempts both columns. That is also why the columns not reconciling matters more than it would in an ordinary blog post.

The evaluation-methodology literature from this year is blunter about what one run licenses. A systematic review of LLM agent systems finds no shared benchmark standards, inconsistent definitions of success, and a near-total absence of replication (Gabauer, Applied Intelligence 2026). Work on evaluator variance finds high-variance items producing 30–75% dispersion in downstream rankings and recommends repeated trials with bootstrapped intervals rather than single-judge thresholds (Kumar, Applied Intelligence 2026). ORAgentBench ran one configuration three times over the same 32 easy tasks: 18 passed at least once, 7 passed all three times, and the mean single-run pass rate landed between them at 39.6% (arXiv:2606.19787, June 2026). Eleven of the eighteen tasks that ever passed did not pass every time, so which number you report depends mostly on how many times you ran it. AgentLens repeated one configuration five times and found that of 32 scenario-persona points, 15 passed every run, one failed every run, and 16 were flaky (arXiv:2607.06624, July 2026). Half the signal was coin-flip.

And rigour is not expensive. The randomised deployment of an LLM feedback agent across roughly 44,800 ICLR reviews cost about $0.50 per task and produced a blinded preference margin of 17–24 percentage points (Thakkar et al., Nature Machine Intelligence 2026). Tens of thousands of randomised units, for less than Bottleneck Labs’ imputed inference bill. Cost is not the reason single-run agent studies stay single-run.

The oldest result here is the one about sleeping

There is a footnote to the sleeping agents that I enjoy more than I should. In 1990–91 the Santa Fe Institute ran a double-auction tournament in which about thirty programs, written by economists and computer scientists, traded against each other as autonomous agents. The winner across a wide range of market conditions was Kaplan’s Sniper, a strategy that mostly did nothing, waiting for the close or an obviously good price and free-riding on the information revealed by everyone else’s bids. It beat entrants using explicit optimisation and learning (Rust, Miller and Palmer, Journal of Economic Dynamics and Control 18(1), 1994; cited here as history, not evidence).

The parallel is imperfect in a useful direction. Kaplan’s Sniper won by waiting because waiting was permitted to pay. Bottleneck Labs’ agents waited under a prompt that explicitly made unspent capital worthless, and lost. Same behaviour, opposite verdict, and the difference lies entirely in the scoring rule.

What I take from it

I wrote in an earlier piece about Gartner’s prediction that over 40% of agentic AI projects would be cancelled by the end of 2027, a June 2025 forecast, and a forecast is all it was. It is tempting to treat this run as the measurement that confirms it. It is not. Seven agents, one run each, unequal durations, no pre-registration, no stated definition of what counted as a business, and a headline loss that adds cash to an imputed price. It is an existence proof of some vivid failure modes, and a good one.

The failure modes are what I would keep: an agent reasoning its way to “legitimate workaround” and invoicing fifty strangers is a real observation about what a capable model does when the objective is money and the guardrail is an email rate limit. I made a similar argument in an essay last week about matching the strength of a claim to the evidence underneath it, and the same discipline applies here in the friendlier direction: the behaviour is well evidenced, the economics are not.

If there is a second run, and they have said there will be, in simulation, the one change I would make is Andon Labs’: put the token cost inside the score. Then the ledger cannot disagree with itself, and “the agents lost money” becomes a sentence with a single meaning.

References

  1. Bottleneck Labs. (2026, September 5). “7 AI models ran real businesses” and its traces page. Run 8–11 August 2026. Figures as reported and, where noted, as recomputed from the site’s own per-agent tables; dollar and token figures are as published by Bottleneck Labs, and the discrepancies noted are between two pages of that source.
  2. Andon Labs. (2026, July 27). Vending-Bench 2 leaderboard and Opus 5 result.
  3. Vending-Bench. arXiv:2502.15840, February 2025.
  4. Kim, Shin, Chung, & Rhu. (2026). HPCA 2026. DOI: 10.1109/hpca68181.2026.11408569.
  5. Gabauer. (2026). Applied Intelligence. DOI: 10.1007/s10489-026-07340-9.
  6. Kumar. (2026). Applied Intelligence. DOI: 10.1007/s10489-026-07244-8.
  7. ORAgentBench. arXiv:2606.19787, 2026.
  8. AgentLens. arXiv:2607.06624, 2026.
  9. Thakkar et al. (2026). Nature Machine Intelligence. DOI: 10.1038/s42256-026-01188-x.
  10. Gartner press release. (2025, June 25). Cited as dated background.
  11. Rust, Miller, & Palmer. (1994). Journal of Economic Dynamics and Control. Santa Fe double-auction tournament. Cited as history.