A customer writes in about a $745 appliance stuck at a carrier hub, fifteen days late. The agent does nine careful tool calls, reads the refund policy correctly, finds that her account segment does not qualify for compensation, and closes the ticket as resolved. The ticket should have been put on hold, because the carrier exception was still open. A grader that reads the transcript sees nine well-formed calls. A grader that reads the database sees one field, status, set to the wrong value.
That example opens a Microsoft and Hugging Face post published on 3 October about ThinkingBox, a sandbox and a 507-task benchmark of stateful business workflows (retail, auto insurance, travel, neobank internal IT, consulting). Its argument comes in three sentences: “A trajectory is a claim. Database state is the evidence. Repetition is the trust test.” I agree with the first half of that and want to look hard at the second. The post’s reliability numbers come with two caveats, one in the post and one in the paper behind it, and the caveats change how much weight the headline can carry.
The part that holds up
Grading the end state is the better design, though it is not new. The 2024 tau-bench paper (Yao et al., arXiv:2406.12045, June 2024) already compared “the database state at the end of a conversation with the annotated goal state,” and it introduced pass^k, the probability that an agent succeeds on all k independent tries. ThinkingBox extends both ideas to more domains and to k = 20. In the paper’s setup, 477 of the 507 tasks are graded on state alone and 30 add a binary rubric for things like required disclosures. Every attempt runs in an isolated session from a clean backend.
The finding that justifies the design is about what transcripts miss. In an ablation of 121,680 valid trials across 12 models, 79,853 failed the executable checks, and 67.24% of those failures still ended cleanly, called a state-changing tool, and reported no final tool error. The agent said it was done and the database disagreed.
Two labels, two values
The post defines its reliability column as “observed 20/20”: the literal count of tasks passed on all 20 attempts, “no estimator, no smoothing.” For Kimi-K3 that is 68 of 507, or 13.41%. The paper’s abstract says Kimi-K3 falls “from 57.37% pass@1 to 17.60% pass^20.” Same model, same benchmark, same 20 runs, two numbers, and the paper’s Table 15 prints both: 17.60% in the pass^20 column and 68 in the 20/20 column.
The difference is an estimator choice. The paper explains that the unbiased tau-bench estimator “is necessarily zero whenever” a task’s success count is below k, so on hard tasks it gives little resolution. It reports a biased plug-in instead, the mean over tasks of (C/20)^20. A task that passed 19 of 20 contributes 0.36 to that mean, not zero. For Kimi-K3 the two definitions are 4.2 points apart. For GPT-5.4 they are 30.62% against 25.25%. For GPT-6 Astra, 46.89% against 45.56%.
For Claude Opus 5 the paper’s pass^20 is 47.53%, which equals 241/507, the observed count, to two decimals. I divided each row’s 20/20 count by 507 and compared it with the pass^20 column: every other model with a nonzero count differs, so this match is the odd one out, and I can’t explain it. It may be coincidence or a transcription slip; the paper doesn’t say. The tau-bench estimator, the plug-in and the raw count are three defensible numbers. The trouble is that a reader comparing the blog’s “observed 20/20” (13.41%) with the paper’s headline “pass^20” (17.60%) meets two numbers for the same idea under different labels.
There is a second mismatch in who was tested. The blog’s tables include Claude Opus 5.5 at 67.16% pass@1 and 241 tasks at 20/20, the same count as Opus 5. The paper’s Table 15 (version 4, 1 October) has 18 models, includes MiniMax-M2.5, and does not include Opus 5.5. I could not find Opus 5.5’s numbers in the paper, so that row is one I can’t check against the paper’s tables. The blog uses it to say that half a point of headline accuracy “bought no additional dependability at all.” Two models landing on 241 may be true. (The blog calls the gap half a point, though its own numbers give two-thirds of one.) It is also a single count with no interval, from a benchmark whose own Opus 5 pass@1 interval runs from 62.78% to 70.20%.
What a 20/20 count mostly measures
This part uses only the paper’s Table 15. Every task falls in one of three bins: never solved in 20 tries, solved every time, or somewhere between. Where a model sits among the three says more than any single number.
Opus 5 has 106 tasks it never solves, 241 it always solves, and 160 in between. GPT-6 Astra splits 147, 231 and 129. Those are bimodal models: a task is either in their repertoire or it isn’t, and when it is, they do it every time. Kimi-K3 splits 31, 68 and 408. About 80% of its tasks land in the middle, and by my arithmetic those 408 tasks account for roughly 77% of its 5,817 successful attempts. Kimi-K3 succeeds almost anywhere and rarely every time.
Those are different engineering problems. For the bimodal model the work is expanding coverage. For the flaky one, the work is variance reduction: retries, validators, a check on terminal state before the commit. A single pass^20 number folds them together, which is how Kimi-K3 can lead on pass@20 (93.89%) and trail on all-20. The paper calls this a gap between discovery and reliability; I would rather read the distribution than the count.
The count is also noisy. The paper’s bootstrap intervals, which resample tasks, put Opus 5’s pass^20 at 47.53% [43.21, 51.84] and GPT-6 Astra’s at 46.89% [42.56, 51.22]. The paper’s line that Claude Opus 5 “leads pass^20 at 47.53%” is true of the point estimates and not of much else. The paper’s top two on pass@1 overlap almost entirely too: Opus 5 at 66.50% [62.78, 70.20] and GPT-5.4 at 65.36% [62.22, 68.50].
Pricing the tail
The post then prices consistency. Cost per dependable task is the estimated cost of the full 20-run campaign divided by the number of tasks that passed 20 of 20. By that measure GPT-5.4 costs $6.80, GPT-6 Astra $7.45 and Claude Opus 5 $13.30, which is how the ranking in the second figure changes.
The post does separate the two prices, and says outright that “the cheapest way to get a right answer is not the cheapest way to get a dependable one.” The paper is more careful than the blog about what the dependable-task price means. Its cost section says the metric “is total campaign cost per observed all-N task, not an estimate of future completion cost,” that both cost metrics “depend on the trial budget,” and that the comparisons “do not certify future reliability.” Prices are the cheapest eligible OpenRouter endpoint on 20 September, undiscounted, and exclude the simulator, the judge and downstream failure. So the second figure ranks 20-run campaigns by what they happened to cost. It does not tell you what a production refund costs.
Tool errors, and a reading I would not stretch
The failure taxonomy has four exclusive buckets, averaged unweighted across models: tool usage 79.9% (range 62.5 to 97.0 across models), wrong state update 10.3%, incomplete user resolution 7.0%, no state-changing action 2.9%. The blog concludes that this “is a retry and error-recovery problem before it is a model problem.” That could be right. The note under the table says the labels are “not unique causal explanations,” and a deterministic signature assigned from a trace cannot tell a model that handles errors badly from a tool surface that is badly designed. The post also says the authors have not measured the lift from any of the mitigations it suggests. I would hold the “not a model problem” sentence at that level of evidence.
Where this stands
Two limits belong next to any number from this work. The authors describe the tasks as “synthetic reconstructions from a non-public source collection” that “are not claimed to represent the distribution of enterprise work,” with fictional customers. The simulated user is an LLM that is cooperative and has a fixed objective, and the paper’s limitations section lists simulator and judge dependence among its own caveats. A model that drops to 17.60% pass^20 (13.41% observed all-20) against a cooperative simulator has told you something about the benchmark’s workflows, and less about how it would behave with an angry customer.
Other 2026 work points the same direction without settling it. Gonzalez-Pumariega et al. (April 2026) analyzed computer-use agents on OSWorld, starting from the premise that an agent that succeeds once may fail on a repeated execution. They examined three factors (execution stochasticity, task ambiguity, behavioral variability) and concluded that reliability depends on how tasks are specified and how agent behavior varies across runs. A noise-floor audit of tool-calling on BFCL (Chen et al., August 2026) found that at temperature 0 reruns flipped only 0.7% to 2.7% of tasks, while semantics-preserving prompt changes moved results 11 to 58 times as much. I take that as a reminder that “repetition” can mean resampling the model, perturbing the prompt, or varying the task, and ThinkingBox does mainly the first.
What I’d take from it: the state-grading harness and the 20-run protocol are worth copying. The code is MIT and the benchmark data is CDLA-Permissive-2.0, so the experiment can be rerun. The single all-20 figure is better read as a shape than a score, and the useful report is the three-bin split per model, with intervals. I wrote about a neighboring problem in the harness-tax study, where gaps of one to seven attempts out of ninety were read as findings. The question is the same one: what is the denominator?
References
- Kundu, T., et al. (2026). The Agent Said It Was Done. The Database Disagreed. Microsoft and Hugging Face blog, 3 October 2026. Source for the opening example, the 121,680-trial ablation, 67.24%, the model tables, “observed 20/20” definitions, the cost figures, the failure-signature shares, and the Opus 5.5 row.
- Li, Z., Ko, Y., Keramati, A., Kundu, T., et al. (2026). One Success Isn’t Reliability: ThinkingBox, a Sandbox and Benchmark for Agents in Stateful Business Workflows. arXiv:2608.19741v4, 1 October 2026 (v1 20 August 2026). Table 15 (pass@1, pass^20, pass@20, 0/20 and 20/20 counts and intervals), Table 23 and Equation 14 (cost), Equations 9 and 10 (estimators), Appendix A (limitations). The cost rankings and three-bin splits in this essay are my arithmetic from Tables 15 and 23.
- Yao, S., Shinn, N., Razavi, P., Narasimhan, K. (2024). tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045, 17 June 2024. Background: end-state comparison and the pass^k metric.
- Gonzalez-Pumariega, G., Agashe, S., Yang, J., Li, A., et al. (2026). On the Reliability of Computer Use Agents. arXiv:2604.17849, 20 April 2026.
- Chen, Y., Qian, P., Wang, S., Peng, C., et al. (2026). Noise Floor Audit for Agent Benchmarks. arXiv:2608.22331, 23 August 2026.
- microsoft/thinkingbox (code) and microsoft/thinkingbox-data (benchmark data, release thinkingbox-bench-v1.0). Code license MIT per the GitHub repository; data license CDLA-Permissive-2.0 per the Hugging Face dataset card, checked 4 October 2026.