Medium Is Not a Unit

One sentence in Anthropic’s Prompting Claude Opus 5.5 guide deserves more attention than the rest of the page: “Effort level names don’t correspond to the same amount of thinking across models”. The guide is undated. It was on Hacker News by this morning, six days after the model’s 22 September launch, and had 194 points and 218 comments at the time of writing. If you have been treating low, medium and high as settings with a fixed meaning, this sentence is the vendor telling you they were never that.

The guide makes two claims about effort that pull in opposite directions. First, Opus 5.5 at medium, its new default, “matches or exceeds Claude Opus 5 at high” on coding and knowledge-work evaluations, so the same answer should now cost you a lower setting. Second, “at a given level, Claude Opus 5.5 tends to think more per turn than Claude Opus 5, especially at xhigh and max,” so leaving your old setting in place buys more thinking than it used to. The label moved in one direction and the tokens behind it moved in the other. The effort documentation settles which one you are actually controlling: “Effort is a behavioral signal, not a strict token budget.”

The Calibrate effort section of Anthropic's Prompting Claude Opus 5.5 guide. It says medium is the default on Opus 5.5 while Opus 5 defaults to high, that effort level names don't correspond to the same amount of thinking across models, and that at a given level Opus 5.5 tends to think more per turn than Opus 5.
Both halves of the problem sit in consecutive paragraphs: the names don't carry across models, and the same name now buys more thinking. Image: Anthropic, "Prompting Claude Opus 5.5," Claude Platform Docs, screenshot of the "Calibrate effort" section taken 28 September 2026, reproduced for commentary.

I don’t read this as a scandal. A behavioral signal is a reasonable design, since a model that thinks less on easy problems and more on hard ones is the whole point of adaptive thinking. But it means the word medium in your config file is a request, and the price of that request is set by whichever model receives it. The migration guide says so outright, if against an older baseline: “The token allocation behind each effort level changes on Claude Opus 5.5 compared to Claude Opus 4.7.” The effort page says the same thing about Sonnet 5.5, whose levels are “recalibrated, so a level doesn’t produce the same amount of thinking as the same level on Claude Sonnet 5.” Each new model resets the scale.

Nine findings, one size

What should you set instead? The guide’s answer is to test several levels against your own evals, which is correct and also a way of saying the guide can’t tell you. What it can tell you, it attributes to Anthropic’s own testing. I count nine such passages (“in Anthropic’s testing” eight times, “in Anthropic’s evaluations” once), and they are the evidence behind nearly every recommendation on the page.

What the nine “in Anthropic’s testing” findings in the Opus 5.5 prompting guide report A grid of the nine findings the Prompting Claude Opus 5.5 guide attributes to Anthropic’s testing or evaluations, checked against four things a reader would need to reuse them: a direction, an effect size, a named benchmark or evaluation set, and a sample size. Eight state a direction; the ninth, on max_tokens, reports a setting that worked rather than a measured effect. One, that a harness reminder roughly halved the share of agentic coding tasks with a long silent stretch, gives an approximate effect size. Several name a broad task category, but none names a benchmark or evaluation set, and none gives a sample size. Source: Anthropic, Prompting Claude Opus 5.5, read 28 September 2026; classification is the writer’s. NINE VENDOR-TESTING CLAIMS · WHAT EACH ONE REPORTS Direction Effect size Named eval Sample size Medium matches or beats Opus 5 at high on repository coding Direction: yes Effect size: no Named eval: no Sample size: no Lowest effort reads dense charts better than Opus 5 at its highest Direction: yes Effect size: no Named eval: no Sample size: no Medium matches or beats Opus 5 at high on coding and knowledge work Direction: yes Effect size: no Named eval: no Sample size: no A max_tokens of 128,000 “has worked well” Direction: no (a setting that worked, not a measured change) Effect size: no Named eval: no Sample size: no A harness reminder roughly halves long silent stretches Direction: yes Effect size: yes Named eval: no Sample size: no An “explore broadly” line gets noticeably more tasks right Direction: yes Effect size: no Named eval: no Sample size: no A time budget lets agent teams finish considerably sooner Direction: yes Effect size: no Named eval: no Sample size: no Dropping “think carefully” makes chat replies start sooner Direction: yes Effect size: no Named eval: no Sample size: no “Treat answers as settled” cuts follow-up thinking Direction: yes Effect size: no Named eval: no Sample size: no 8 of 9 1 of 9 0 of 9 0 of 9 Filled = reported, open = not. “Roughly halved” counted as an effect size. Writer’s classification of the guide as read 28 Sep 2026.
Eight of the guide’s nine in-house findings tell you which way things moved. One tells you roughly how far. None tells you on what, or how many.

Eight of the nine say which way something moved. Only one says roughly how far: a harness reminder “roughly halved the share of tasks with a long silent stretch, with no measurable change in cost.” The rest are “noticeably more,” “considerably sooner,” “no clear decline,” “a small fraction.” Several name a broad category of task (agentic coding, multi-app automation, research tasks for small agent teams), but none names a benchmark, a run date or a count. The ninth is advice about a max_tokens setting that “has worked well,” which is not a measurement at all. I don’t doubt any of the nine is true in the direction stated. You just can’t size any of them against your own workload, and sizing is the point of the exercise the guide recommends.

Anthropic has the numbers. A sibling page on the same docs site, Optimizing for cost and intelligence, reads like an appendix the prompting guide forgot to link. It gives run dates, problem counts, how many runs were averaged, 95% intervals, and in one experiment a noise margin “that Anthropic set before the runs.” On a 478-problem subset of SWE-bench Pro, run on 19 and 20 September, it prices each effort level against high:

What each effort level costs and buys on SWE-bench Pro, Claude Opus 5.5, measured against high Four points plotting score change against cost, both relative to the high effort level, on a subset of SWE-bench Pro. Low scored about 8 points lower for about a third of the cost. Medium, the default, scored about 2.5 points lower for about 70 percent of the cost. High is the reference. Xhigh scored about 1.4 points higher for 2.5 times the cost. The cost axis is logarithmic. Source: Anthropic, Optimizing for cost and intelligence, runs of 19 to 20 September 2026 on a 478-problem subset; figures are Anthropic’s approximations. OPUS 5.5 · SWE-BENCH PRO SUBSET · EACH LEVEL AGAINST HIGH 0 pts -4 pts -8 pts 0.25× 0.5× 1× 2× 3× cost per task relative to high (log scale) low: about 8 points lower, about a third of the cost — Anthropic, runs of 19–20 Sep 2026 medium: about 2.5 points lower, about 70% of the cost — Anthropic, runs of 19–20 Sep 2026 high: reference — Anthropic, runs of 19–20 Sep 2026 xhigh: about 1.4 points higher, 2.5 times the cost — Anthropic, runs of 19–20 Sep 2026 low medium (default) high xhigh Anthropic’s rounded figures (“about”), 478 problems; low, medium and high average two runs, xhigh is one run.
On long-horizon coding, the default costs something. Medium gives back about two and a half points against high on Anthropic’s own subset, and it is still the level the guide tells you to start from.

The numbers are useful, and they don’t quite match the guide’s tone. medium gave back about 2.5 points against Opus 5.5’s own high for about 70% of the cost. low gave back about 8 points for about a third. xhigh added about 1.4 points, from a single run, for 2.5 times the cost of high. The same page is candid that this is where effort matters most (“long-horizon coding is where effort genuinely buys accuracy”), while on four research and knowledge-work benchmarks, measured on Fable 5, medium matched the default’s accuracy at about 70% to 87% of its cost. Different model, different work, so the comparison is loose, but the direction is plain: the exchange rate between a label and a result depends on the task. A two-and-a-half-point loss on one and no measurable loss on the other is the same word doing two different jobs.

The comparison the guide’s central effort claim rests on, Opus 5.5 at medium against Opus 5 at a stated high, never appears as a number in Anthropic’s own prose. The cost page never sets the two side by side: its SWE-bench Pro charts carry Opus 5 at its default (high) on a 482-problem subset in August and Opus 5.5 on a 478-problem version in September, in separate charts. The launch post’s benchmark text and the system card’s text mostly compare against Opus 5 at max instead. The one exception, FrontierCode, sets medium against Opus 5 at its unstated “best reasoning effort” (54.6% to 53.4%), and the card’s charts plot every level on log axes without printing the values. On CursorBench 4.0, Opus 5.5 at medium scored 52.5% at about $3 a task against 46.6% at $11.95 for Opus 5 at max (Cursor measured the scores; the Opus 5.5 cost per task is Anthropic’s estimate from Cursor’s token counts). That is a stronger claim than the guide’s, on one benchmark, and a different one.

The like-for-like numbers come from customers quoted in the launch post. Factory says Opus 5.5 at medium “matched Opus 5 on high effort, while using 20 to 25% fewer output tokens.” Deloitte says that even at its lowest setting it “caught 72% of known bugs in our code reviews to Opus 5’s 56% at high effort.” Those are testimonials the vendor chose, on the customers’ own workloads, with no counts attached. The guide’s claim may well hold. The evidence printed for it is a customer quote.

The sibling page also shows how noisy these comparisons get. On Chartography, a chart-reading benchmark, Opus 5.5 at low scored 68.7 against 49 for Opus 5 at low. That is low against low, run with tools, so it is consistent with the guide’s stronger claim (the new model’s lowest setting beating the old one’s highest, without tools) rather than a test of it. The footnote adds that “run-to-run spreads were up to 10 points.” A gap of nearly 20 points survives that. Plenty of the gaps you will measure in your own sweep won’t, which is roughly what I found when I went through the HarnessTax numbers last week.

Across vendors, it’s worse

Within one vendor, the labels at least keep their order. Across vendors, they don’t have a common scale at all. The best recent measurement I found is a side result in an arXiv preprint on draft-verify-revise pipelines (10 September), which ran six models from three providers across 21 effort configurations. Mean reasoning tokens never fell as the effort level rose, which is reassuring. But the peak each model reached “varied roughly sevenfold across models”: 1,076 reasoning tokens per trial for Claude Opus 4.6 at max, 2,613 for GPT-5-mini at high, and 7,611 for GPT-5.2 at xhigh. By that measure Anthropic’s top setting spent less than GPT-5-mini did at high.

Two caveats keep this from being a league table. The dataset is small (10 base examples in three conditions), and the providers don’t even report the same quantity: OpenAI and Google return a reasoning-token count, while Anthropic returns one output total, so the author derived Claude’s reasoning tokens by subtracting the measured length of the visible reply. Opus 4.6 also helped write the test material, which the paper flags. The paper’s data were also collected in February, several model generations back. Still, the direction is clear enough. high means one thing at one company and something else at another. It also changes meaning within a company every time a new model ships.

The instruction that used to be the thinking

The guide’s other thread is about prompts that did the thinking for the model. It tells developers moving off Opus 5 with thinking disabled to “remove instructions that stood in for thinking.” If a prompt asks the model to write out its reasoning in the reply as a substitute, the guide says to delete it and read summarized thinking blocks instead, and warns that such a prompt “can be declined with the reasoning_extraction refusal category.” A chat system prompt that says “think carefully before answering” should probably go too, since in Anthropic’s testing removing it made replies start sooner “with no clear decline in the quality of the reply.”

Four years ago that instruction was the thinking. In May 2022, Kojima and colleagues showed in “Large Language Models are Zero-Shot Reasoners” that appending “Let’s think step by step” to a question moved text-davinci-002 from 17.7% to 78.7% on the MultiArith arithmetic set. A single sentence of prompt did what a feature of the model now does. The guide now suggests removing its descendants, the “think carefully” lines in chat system prompts, and says a prompt that pushes the model to reproduce its reasoning in the reply can be declined. The prompting habits of 2022 are the ones being cleaned out.

Much of the Hacker News thread was about that churn. One commenter, mathisfun123, wrote that the advice “changes like every 3 months.” Another, buckle8017, reduced the whole guide to “Medium is about as good as 5.0 was at high,” which is a fair summary of a claim the guide never puts a number on.

What a unit would look like

When I worked through the Opus 5.5 price cut last week, I held token counts fixed and let the rate card do the work. Effort is where that assumption breaks. The rate card sets the price per token, and effort, through whatever the model decides a label means this generation, sets how many tokens you buy. A model with a 20% lower per-token price can still cost you more if it thinks more per turn at your old setting. A couple of Hacker News commenters reported longer answers or usage limits running out sooner, without tying either to effort; those are anecdotes, not invoices.

A unit would be something like a published table: for each model and each level, the median and 90th-percentile thinking tokens per turn on a named task set, with the run date. Anthropic’s own cost page already reports a per-turn output-token distribution for Opus 5.5 at its default, dated 19 September, with the longest turn at about 61,000 tokens. That is a start, but it covers one level and doesn’t split out thinking. Until the prompting guide carries one, the only calibrated scale for medium is the sweep you run yourself, on your own traffic, again, every time the model string changes.

References

  1. Anthropic (2026). Prompting Claude Opus 5.5. Claude Platform Docs, undated page, read 28 September 2026.
  2. Anthropic (2026). Introducing Claude Opus 5.5. 22 September 2026.
  3. Anthropic (2026). Effort. Claude Platform Docs, undated page, read 28 September 2026.
  4. Anthropic (2026). Migrating to Claude Opus 5.5. Claude Platform Docs, undated page, read 28 September 2026.
  5. Anthropic (2026). Optimizing for cost and intelligence. Claude Platform Docs, undated page, read 28 September 2026; SWE-bench Pro runs of 19–20 September 2026, Chartography runs of 6 and 9 August and 20 September 2026.
  6. Anthropic (2026). Claude Opus 5.5 System Card. 22 September 2026, section 8.8 (CursorBench 4.0).
  7. Ekekezie, O. I. (2026). Can LLMs in Draft-Verify-Revise Pipelines Resolve Deictic Ambiguity? arXiv:2609.12162, 10 September 2026, Supplementary Figure S6 and Methods.
  8. Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., & Iwasawa, Y. (2022). Large Language Models are Zero-Shot Reasoners. arXiv:2205.11916, 24 May 2022, Table 1.
  9. Hacker News (2026). Prompting Claude Opus 5.5. Thread submitted 28 September 2026, 194 points and 218 comments when read; comments by mathisfun123 and buckle8017.