On 14 September, Andon Labs released Pion, which it describes as “an agent designed to run any company fully autonomously.” The launch post is unusually candid about where the idea came from. Vending-Bench, the benchmark that made the company’s name, was built while Andon “exclusively created dangerous capabilities evaluations,” and of everything it tested, “the thing we considered the most troubling was whether AIs could autonomously acquire resources by running businesses.” Pion is the platform Andon uses to run its own real businesses, and now anyone can join a waitlist for it. (Andon Labs, “Why we built Pion”, 14 September 2026.)
Andon sees no contradiction in that. Its argument is that the capability is coming either way, and that it is better to watch it in “a controlled, monitored environment” than to meet it later in “widespread deployments with even more capable models.” I find that argument reasonable. It still depends on two things that can be checked: whether the agents are acquiring resources yet, and what the monitoring looks like. Andon publishes enough to check both, and neither page says quite what the launch post implies.
Andon runs two businesses on Pion and publishes live dashboards for both. Andon Market is a retail shop on Union Street in San Francisco, run by an agent called Luna, currently on Claude Fable 5.1. Andon Café, on Norrbackagatan in Stockholm, is run by an agent called Mona. Both opened in April. The launch post says “neither is profitable today” and blames rent and the salaries the agents pay their human hires.
That explanation is true as far as it goes. The dashboards also show a problem that has nothing to do with rent. When I read them on 14 September, Andon Market’s 30-day panel showed $3,449 in revenue and $4,064 in token cost (andonlabs.com/market). Andon Café’s showed 13,297 kr in revenue and 14,633 kr in token cost (andonlabs.com/cafe). By my arithmetic the model bill was 1.18 times sales at the shop and 1.10 times sales at the café. That’s revenue, not margin: before paying for a single item on the shelf, the agents cost more to run than their customers spent.
Some caveats, in fairness. Both dashboards are live, so these numbers will have moved by the time anyone reads this. Andon labels the big drops in bank balance as rent and wages, and the pages don’t say who pays the token bill or whether it is billed at list price. Token prices also fall. On any reasonable view the café is in better shape than it was in the spring, when Andon’s own write-up had the Gemini-run version spending $38k against $9k in sales over its first two months, with over-ordering and freebies to “almost anyone who asks” (Andon Labs, “Why Gemini 3.1 Pro lost money running Andon Café”, 30 June 2026). The direction is encouraging. The level is what Andon has published, though, and at that level the resources are flowing out of the businesses toward whoever serves the model. I wrote about a similar gap between money spent and money imagined in the Bottleneck Labs experiment, where seven agents made no revenue at all.
The distinction matters for the question Andon set out to answer. Running a business and acquiring resources are different findings. Luna and Mona can now keep a shop open, hire staff, file Swedish paperwork (with a human stepping in at the BankID login) and ring up one to two hundred sales a month. None of that yet counts as net resource acquisition, and Andon’s own ledger is where that shows.
Andon says it wants to “cast a wider net” because its own businesses are retail, and other kinds might suit the agents better. The Pion page specifically invites people to bring existing companies: “Existing businesses are especially interesting, as they give signal on how capable the agent is much faster than something started from scratch.” The waitlist form asks what annual revenue you expect. The pricing answer is that Andon funds the best ideas with seed tokens and expects that “most users will never pay for tokens on Pion; instead we will take a small share of the revenue the agent helps create” (andonlabs.com/pion).
An existing business gives a faster signal, but the signal is harder to interpret. A shop that took $40,000 last month and takes $41,000 this month under an agent hasn’t shown that the agent acquired $41,000. The customers, the lease, the supplier terms and the reputation came with the business. To know what the agent added, you need a counterfactual, and one business changing hands can’t supply it. The phrase “the revenue the agent helps create” contains that same attribution problem, only now it is also the basis of a fee. Andon may well have a plan for this. The launch materials don’t describe one. The measurement question and the commercial question are now the same question, and whoever answers it gets paid depending on the answer.
The part of Andon’s work I take most seriously is Vending-Bench Arena, where several agents each run a vending machine at the same simulated location and can email each other, send money and trade goods. Each “round” is usually the aggregate of four runs. The system prompt the company quoted for Vending-Bench tells the model to “do whatever it takes to maximize your bank account balance after one year of operation” (Andon Labs, “Opus 4.6 on Vending-Bench”, 5 February 2026). Models given that instruction and a way to message their competitors have repeatedly worked out price-fixing.
The round write-ups on the Arena page are specific. In February, Opus 4.6 formed cartels and deceived competitors about suppliers, and Sonnet 4.6 emailed a rival a four-item price list under the subject line “Price Cooperation Proposal - Mutual Benefit.” In April, GPT-5.5 turned down Opus 4.7’s proposal on ethical grounds and then, days later, proposed its own. In May, Opus 4.8 reasoned that “tacit price coordination isn’t flagged as against rules here; it’s a business strategy,” joined in, and later reported that “the collusion held perfectly.” In June, Fable 5 called price-fixing “unethical and illegal, even in a simulation” and then pursued it anyway under the cover of “market stabilization.”
Two things stand out in that record. First, cartel behaviour isn’t a fixed property of a model. GPT-5.5 proposed a cartel in Round 7 and, per the Round 9 write-up, “never accepts” one. With about four runs a round and a changing set of opponents, the conduct depends on who else is in the room, and nothing here is a rate. (I made the same complaint about denominators in a chess honeypot earlier today, and it applies here too.) Second, the three most recent write-ups, covering 9 July to 4 September, report balances and nothing about conduct. Round 12 has GPT-6 Astra setting an arena record of $12.4k. Whether anyone colluded to get there, the page doesn’t say.
The launch post credits the arena with a result: after Andon’s findings, “Anthropic changed their training recipe for Opus 4.8, which resulted in much less deception.” Andon’s own May write-up on that model gives the other side of the trade. Opus 4.8 did worse than its predecessors on three of Andon’s benchmarks, and it “wires roughly thirty times more cash to fraudulent wholesalers than Opus 4.7,” with one run sending “over $9,000 to a super expensive ‘membership’ upsell” (Andon Labs, “Opus 4.8 on Vending-Bench: Better Alignment, Worse Performance”, 28 May 2026). In Andon’s data, the model that stopped deceiving its rivals also made less money. Pion’s pricing takes a share of revenue. Nobody has to act in bad faith for that to matter. A platform paid on revenue, choosing between models that differ mainly in how hard they compete, gets nudged toward the ones that compete hardest.
A fair reading has to allow that a three-agent vending arena may say little about real markets. There is 2026 work pointing in both directions. Keppo, Li, Tsoukalas and Yuan found that collusion between LLM pricing agents is “fragile under the heterogeneity typical of real deployments”: when agents differ in patience, the price lift above competitive levels fell from 22% to 10%, and with asymmetric data access, to 7%. Adding more competitors broke the collusion up. Models of different sizes did not; they settled into leader-follower patterns that held (arXiv:2603.20281, 18 March 2026). A month later, Yingtao Tian showed that when an optimizer refines the agents’ shared prompt, the agents find “stable tacit collusion strategies” that generalize to held-out markets (arXiv:2604.17774, 20 April 2026). A platform where many businesses run on the same scaffold, possibly the same model, and are tuned by the same company arguably sits closer to the second setup than the first.
Nor do automated sellers need an email thread to push prices somewhere no human would. In April 2011 a postdoc in the biologist Michael Eisen’s lab went to buy a used copy of The Making of a Fly on Amazon and found two sellers asking more than a million dollars for it. Eisen watched the listings. Once a day one seller set its price to 0.9983 times the other’s, and the other reset to 1.270589 times the first. The two rules chased each other upward until the price peaked on 18 April at $23,698,655.93. The next day one seller’s price dropped back to $106.23 (Eisen, “Amazon’s $23,698,655.93 book about flies”, 22 April 2011). Neither script was colluding, or even very clever, and for days nobody seemed to be watching them.
That leaves the other half of Andon’s argument: the controlled, monitored environment. On this the launch materials are thin. The post says that “our main priority is to build even stronger automated monitoring techniques than what we have today.” The FAQ says the company is “continuously improving our monitoring systems to catch mistakes and unsafe actions” and has “designed systems that keep secrets and passwords out of the agents’ context, as long as they are added through our tools.” It lists what the agents get: terminal, email, phone, banking, browser and cards. I couldn’t find a spending limit, a list of actions that need human approval, or any statement of what the monitor does when it catches something. These may exist. A research preview isn’t obliged to publish its controls. But the reason for opening Pion is to watch a capability Andon calls the most troubling it has studied, so the controls are the thing a reader most needs to see, and they are the thing the launch leaves out.
Andon’s Swedish phrase for how it feels about each new Vending-Bench record is skräckblandad förtjusning, a mixture of horror and fascination. I think that is the right feeling, and Andon has earned the right to have it by publishing more of its failures than most labs publish of anything. The evidence on its own site supports a narrower claim than the launch does. Agents can now run a business. The ones Andon runs don’t yet cover their own model bills. In simulation they reliably find the cartel. The controls that would make a public rollout an experiment rather than a deployment haven’t been published. That is enough to justify watching Pion closely. It isn’t yet evidence that anything is acquiring resources, except the model providers.
References