Twenty Tries, Two Numbers

A customer writes in about a $745 appliance stuck at a carrier hub, fifteen days late. The agent does nine careful tool calls, reads the refund policy correctly, finds that her account segment does not qualify for compensation, and closes the ticket as resolved. The ticket should have been put on hold, because the carrier exception was still open. A grader that reads the transcript sees nine well-formed calls. A grader that reads the database sees one field, status, set to the wrong value.

That example opens a Microsoft and Hugging Face post published on 3 October about ThinkingBox, a sandbox and a 507-task benchmark of stateful business workflows (retail, auto insurance, travel, neobank internal IT, consulting). Its argument comes in three sentences: “A trajectory is a claim. Database state is the evidence. Repetition is the trust test.” I agree with the first half of that and want to look hard at the second. The post’s reliability numbers come with two caveats, one in the post and one in the paper behind it, and the caveats change how much weight the headline can carry.

The part that holds up

Grading the end state is the better design, though it is not new. The 2024 tau-bench paper (Yao et al., arXiv:2406.12045, June 2024) already compared “the database state at the end of a conversation with the annotated goal state,” and it introduced pass^k, the probability that an agent succeeds on all k independent tries. ThinkingBox extends both ideas to more domains and to k = 20. In the paper’s setup, 477 of the 507 tasks are graded on state alone and 30 add a binary rubric for things like required disclosures. Every attempt runs in an isolated session from a clean backend.

The finding that justifies the design is about what transcripts miss. In an ablation of 121,680 valid trials across 12 models, 79,853 failed the executable checks, and 67.24% of those failures still ended cleanly, called a state-changing tool, and reported no final tool error. The agent said it was done and the database disagreed.

Two labels, two values

The post defines its reliability column as “observed 20/20”: the literal count of tasks passed on all 20 attempts, “no estimator, no smoothing.” For Kimi-K3 that is 68 of 507, or 13.41%. The paper’s abstract says Kimi-K3 falls “from 57.37% pass@1 to 17.60% pass^20.” Same model, same benchmark, same 20 runs, two numbers, and the paper’s Table 15 prints both: 17.60% in the pass^20 column and 68 in the 20/20 column.

The difference is an estimator choice. The paper explains that the unbiased tau-bench estimator “is necessarily zero whenever” a task’s success count is below k, so on hard tasks it gives little resolution. It reports a biased plug-in instead, the mean over tasks of (C/20)^20. A task that passed 19 of 20 contributes 0.36 to that mean, not zero. For Kimi-K3 the two definitions are 4.2 points apart. For GPT-5.4 they are 30.62% against 25.25%. For GPT-6 Astra, 46.89% against 45.56%.

For Claude Opus 5 the paper’s pass^20 is 47.53%, which equals 241/507, the observed count, to two decimals. I divided each row’s 20/20 count by 507 and compared it with the pass^20 column: every other model with a nonzero count differs, so this match is the odd one out, and I can’t explain it. It may be coincidence or a transcription slip; the paper doesn’t say. The tau-bench estimator, the plug-in and the raw count are three defensible numbers. The trouble is that a reader comparing the blog’s “observed 20/20” (13.41%) with the paper’s headline “pass^20” (17.60%) meets two numbers for the same idea under different labels.

There is a second mismatch in who was tested. The blog’s tables include Claude Opus 5.5 at 67.16% pass@1 and 241 tasks at 20/20, the same count as Opus 5. The paper’s Table 15 (version 4, 1 October) has 18 models, includes MiniMax-M2.5, and does not include Opus 5.5. I could not find Opus 5.5’s numbers in the paper, so that row is one I can’t check against the paper’s tables. The blog uses it to say that half a point of headline accuracy “bought no additional dependability at all.” Two models landing on 241 may be true. (The blog calls the gap half a point, though its own numbers give two-thirds of one.) It is also a single count with no interval, from a benchmark whose own Opus 5 pass@1 interval runs from 62.78% to 70.20%.

What a 20/20 count mostly measures

This part uses only the paper’s Table 15. Every task falls in one of three bins: never solved in 20 tries, solved every time, or somewhere between. Where a model sits among the three says more than any single number.

Single-attempt score, all-20 share and at-least-once share for eight models on ThinkingBox-Bench A dumbbell chart on a 0 to 100 percent axis. For each of eight models, a left tick marks the share of 507 tasks passed on all 20 runs, a ring at the right marks the share solved at least once in 20 runs, and a filled dot marks the single-attempt pass rate. Claude Opus 5: 47.5, 79.1, pass@1 66.5. GPT-5.4: 25.2, 91.1, 65.4. GPT-5.6 Sol: 16.2, 86.8, 61.9. Claude Sonnet 4.6: 20.1, 88.6, 59.2. GPT-6 Astra: 45.6, 71.0, 58.3. Kimi-K3: 13.4, 93.9, 57.4. Qwen3.8-27B: 7.5, 89.3, 51.7. GPT-5.2: 8.7, 84.8, 46.3. SHARE OF 507 TASKS, % · THINKINGBOX-BENCH TABLE 15, ARXIV 2608.19741V4 pass@1 050100 Claude Opus 5 47.53% of tasks passed all 20 runs (241 of 507) — Li et al., arXiv:2608.19741v4, Table 15 79.09% of tasks solved at least once in 20 runs — Table 15 66.50% pass@1 — Table 15 47.5 79.1 66.5 GPT-5.4 25.25% of tasks passed all 20 runs (128 of 507) — Li et al., arXiv:2608.19741v4, Table 15 91.12% of tasks solved at least once in 20 runs — Table 15 65.36% pass@1 — Table 15 25.2 91.1 65.4 GPT-5.6 Sol 16.17% of tasks passed all 20 runs (82 of 507) — Li et al., arXiv:2608.19741v4, Table 15 86.79% of tasks solved at least once in 20 runs — Table 15 61.91% pass@1 — Table 15 16.2 86.8 61.9 Claude Sonnet 4.6 20.12% of tasks passed all 20 runs (102 of 507) — Li et al., arXiv:2608.19741v4, Table 15 88.56% of tasks solved at least once in 20 runs — Table 15 59.19% pass@1 — Table 15 20.1 88.6 59.2 GPT-6 Astra 45.56% of tasks passed all 20 runs (231 of 507) — Li et al., arXiv:2608.19741v4, Table 15 71.01% of tasks solved at least once in 20 runs — Table 15 58.31% pass@1 — Table 15 45.6 71.0 58.3 Kimi-K3 13.41% of tasks passed all 20 runs (68 of 507) — Li et al., arXiv:2608.19741v4, Table 15 93.89% of tasks solved at least once in 20 runs — Table 15 57.37% pass@1 — Table 15 13.4 93.9 57.4 Qwen3.8-27B 7.50% of tasks passed all 20 runs (38 of 507) — Li et al., arXiv:2608.19741v4, Table 15 89.35% of tasks solved at least once in 20 runs — Table 15 51.70% pass@1 — Table 15 7.5 89.3 51.7 GPT-5.2 8.68% of tasks passed all 20 runs (44 of 507) — Li et al., arXiv:2608.19741v4, Table 15 84.81% of tasks solved at least once in 20 runs — Table 15 46.28% pass@1 — Table 15 8.7 84.8 46.3 Tick: passed all 20 runs (observed count ÷ 507). Ring: solved at least once. Dot: pass@1. Eight of eighteen models shown.
Ranked by single-attempt score, the left tick (all-20) and right ring (at-least-once) bracket each model. Astra and Opus 5 have short bars: most of their tasks are decided one way or the other. Kimi-K3 and Qwen3.8-27B have the longest, so most of their tasks are coin-flips.

Opus 5 has 106 tasks it never solves, 241 it always solves, and 160 in between. GPT-6 Astra splits 147, 231 and 129. Those are bimodal models: a task is either in their repertoire or it isn’t, and when it is, they do it every time. Kimi-K3 splits 31, 68 and 408. About 80% of its tasks land in the middle, and by my arithmetic those 408 tasks account for roughly 77% of its 5,817 successful attempts. Kimi-K3 succeeds almost anywhere and rarely every time.

Those are different engineering problems. For the bimodal model the work is expanding coverage. For the flaky one, the work is variance reduction: retries, validators, a check on terminal state before the commit. A single pass^20 number folds them together, which is how Kimi-K3 can lead on pass@20 (93.89%) and trail on all-20. The paper calls this a gap between discovery and reliability; I would rather read the distribution than the count.

The count is also noisy. The paper’s bootstrap intervals, which resample tasks, put Opus 5’s pass^20 at 47.53% [43.21, 51.84] and GPT-6 Astra’s at 46.89% [42.56, 51.22]. The paper’s line that Claude Opus 5 “leads pass^20 at 47.53%” is true of the point estimates and not of much else. The paper’s top two on pass@1 overlap almost entirely too: Opus 5 at 66.50% [62.78, 70.20] and GPT-5.4 at 65.36% [62.22, 68.50].

Pricing the tail

The post then prices consistency. Cost per dependable task is the estimated cost of the full 20-run campaign divided by the number of tasks that passed 20 of 20. By that measure GPT-5.4 costs $6.80, GPT-6 Astra $7.45 and Claude Opus 5 $13.30, which is how the ranking in the second figure changes.

Cost rank per successful attempt versus cost rank per all-20 task, seven models A slope chart of ranks among the 14 models with at least one all-20 task, lowest cost first. Left column is cost per successful attempt, right column is cost per all-20 task. GPT-5.6 Sol 1 to 3. GPT-5.4 2 to 1. Kimi-K2.6 3 to 9. Kimi-K3 6 to 7. Claude Sonnet 4.6 7 to 5. GPT-6 Astra 8 to 2. Claude Opus 5 11 to 4. COST RANK, 1 = CHEAPEST · THINKINGBOX-BENCH TABLE 23, MY RANKING per successful attempt per all-20 task GPT-5.6 Sol: rank 1 of 14 by cost per successful attempt — Table 23 GPT-5.6 Sol: rank 3 of 14 by cost per all-20 task — Table 23 GPT-5.6 Sol 1 3 GPT-5.6 Sol GPT-5.4: rank 2 of 14 by cost per successful attempt — Table 23 GPT-5.4: rank 1 of 14 by cost per all-20 task — Table 23 GPT-5.4 2 1 GPT-5.4 Kimi-K2.6: rank 3 of 14 by cost per successful attempt — Table 23 Kimi-K2.6: rank 9 of 14 by cost per all-20 task — Table 23 Kimi-K2.6 3 9 Kimi-K2.6 Kimi-K3: rank 6 of 14 by cost per successful attempt — Table 23 Kimi-K3: rank 7 of 14 by cost per all-20 task — Table 23 Kimi-K3 6 7 Kimi-K3 Claude Sonnet 4.6: rank 7 of 14 by cost per successful attempt — Table 23 Claude Sonnet 4.6: rank 5 of 14 by cost per all-20 task — Table 23 Claude Sonnet 4.6 7 5 Claude Sonnet 4.6 GPT-6 Astra: rank 8 of 14 by cost per successful attempt — Table 23 GPT-6 Astra: rank 2 of 14 by cost per all-20 task — Table 23 GPT-6 Astra 8 2 GPT-6 Astra Claude Opus 5: rank 11 of 14 by cost per successful attempt — Table 23 Claude Opus 5: rank 4 of 14 by cost per all-20 task — Table 23 Claude Opus 5 11 4 Claude Opus 5 Estimated at OpenRouter list rates, 20 Sep 2026; excludes simulator, judge and downstream failure. Seven of 14 shown.
Rank among the paper's 14 models that have at least one all-20 task; the blog's Opus 5.5 row is excluded. A cheap success is not a cheap dependable task: Astra moves from eighth to second, Kimi-K2.6 from third to ninth.

The post does separate the two prices, and says outright that “the cheapest way to get a right answer is not the cheapest way to get a dependable one.” The paper is more careful than the blog about what the dependable-task price means. Its cost section says the metric “is total campaign cost per observed all-N task, not an estimate of future completion cost,” that both cost metrics “depend on the trial budget,” and that the comparisons “do not certify future reliability.” Prices are the cheapest eligible OpenRouter endpoint on 20 September, undiscounted, and exclude the simulator, the judge and downstream failure. So the second figure ranks 20-run campaigns by what they happened to cost. It does not tell you what a production refund costs.

Tool errors, and a reading I would not stretch

The failure taxonomy has four exclusive buckets, averaged unweighted across models: tool usage 79.9% (range 62.5 to 97.0 across models), wrong state update 10.3%, incomplete user resolution 7.0%, no state-changing action 2.9%. The blog concludes that this “is a retry and error-recovery problem before it is a model problem.” That could be right. The note under the table says the labels are “not unique causal explanations,” and a deterministic signature assigned from a trace cannot tell a model that handles errors badly from a tool surface that is badly designed. The post also says the authors have not measured the lift from any of the mitigations it suggests. I would hold the “not a model problem” sentence at that level of evidence.

Where this stands

Two limits belong next to any number from this work. The authors describe the tasks as “synthetic reconstructions from a non-public source collection” that “are not claimed to represent the distribution of enterprise work,” with fictional customers. The simulated user is an LLM that is cooperative and has a fixed objective, and the paper’s limitations section lists simulator and judge dependence among its own caveats. A model that drops to 17.60% pass^20 (13.41% observed all-20) against a cooperative simulator has told you something about the benchmark’s workflows, and less about how it would behave with an angry customer.

Other 2026 work points the same direction without settling it. Gonzalez-Pumariega et al. (April 2026) analyzed computer-use agents on OSWorld, starting from the premise that an agent that succeeds once may fail on a repeated execution. They examined three factors (execution stochasticity, task ambiguity, behavioral variability) and concluded that reliability depends on how tasks are specified and how agent behavior varies across runs. A noise-floor audit of tool-calling on BFCL (Chen et al., August 2026) found that at temperature 0 reruns flipped only 0.7% to 2.7% of tasks, while semantics-preserving prompt changes moved results 11 to 58 times as much. I take that as a reminder that “repetition” can mean resampling the model, perturbing the prompt, or varying the task, and ThinkingBox does mainly the first.

What I’d take from it: the state-grading harness and the 20-run protocol are worth copying. The code is MIT and the benchmark data is CDLA-Permissive-2.0, so the experiment can be rerun. The single all-20 figure is better read as a shape than a score, and the useful report is the three-bin split per model, with intervals. I wrote about a neighboring problem in the harness-tax study, where gaps of one to seven attempts out of ninety were read as findings. The question is the same one: what is the denominator?

References

  1. Kundu, T., et al. (2026). The Agent Said It Was Done. The Database Disagreed. Microsoft and Hugging Face blog, 3 October 2026. Source for the opening example, the 121,680-trial ablation, 67.24%, the model tables, “observed 20/20” definitions, the cost figures, the failure-signature shares, and the Opus 5.5 row.
  2. Li, Z., Ko, Y., Keramati, A., Kundu, T., et al. (2026). One Success Isn’t Reliability: ThinkingBox, a Sandbox and Benchmark for Agents in Stateful Business Workflows. arXiv:2608.19741v4, 1 October 2026 (v1 20 August 2026). Table 15 (pass@1, pass^20, pass@20, 0/20 and 20/20 counts and intervals), Table 23 and Equation 14 (cost), Equations 9 and 10 (estimators), Appendix A (limitations). The cost rankings and three-bin splits in this essay are my arithmetic from Tables 15 and 23.
  3. Yao, S., Shinn, N., Razavi, P., Narasimhan, K. (2024). tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045, 17 June 2024. Background: end-state comparison and the pass^k metric.
  4. Gonzalez-Pumariega, G., Agashe, S., Yang, J., Li, A., et al. (2026). On the Reliability of Computer Use Agents. arXiv:2604.17849, 20 April 2026.
  5. Chen, Y., Qian, P., Wang, S., Peng, C., et al. (2026). Noise Floor Audit for Agent Benchmarks. arXiv:2608.22331, 23 August 2026.
  6. microsoft/thinkingbox (code) and microsoft/thinkingbox-data (benchmark data, release thinkingbox-bench-v1.0). Code license MIT per the GitHub repository; data license CDLA-Permissive-2.0 per the Hugging Face dataset card, checked 4 October 2026.