437x Against a 2023 Baseline: What the DeepSeek Cache Number Does and Doesn't Show

A post on insufferable.dev, dated 29 September, argues that the Western labs are now adopting Chinese labs’ work rather than distilling it. Its evidence has two parts: a 437x drop in KV-cache size from DeepSeek, and two cache-read price cuts at Anthropic and OpenAI that it calls “the proof of the adoption.” The post has no personal byline, and it links to nothing that backs these figures: not the paper, not either price page.

I checked both parts. The first is accurate and narrower than it sounds. The second is arithmetically correct and doesn’t prove what it is offered to prove.

Where 437 comes from

The number is in DeepSeek’s own paper on V4.1-Flash (arXiv 2609.19969, submitted 17 September). Figure 1(b) charts “global KV cache per token” in bytes for four models: V1 at 389,120, V3.2 at 48,068, V4-Flash at 3,514, and V4.1-Flash at 890. The caption states the reductions against V4-Flash and V1 as roughly 4-fold and 437-fold. Divide 389,120 by 890 and you get 437.2. The blog’s arithmetic is fine.

Bar chart titled Global KV Cache Per Token (Bytes). DeepSeek-V1 (2023.11) is a long grey bar at 389,120; V3.2 (2025.12) is much shorter at 48,068; V4-Flash (2026.04) is a sliver at 3,514; V4.1-Flash (2026.09) is barely visible at 890. Curved arrows between rows are labelled 8.1x, 13.7x and 3.9x smaller.
The chart behind the headline. On a linear axis the newest model's bar is too thin to see. The step ratios are the paper's rounded figures. Image: DeepSeek-AI, "DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression," arXiv:2609.19969, Figure 1(b). Reproduced for commentary; the paper carries arXiv's non-exclusive distribution license, not an open one.

Three things about the baseline matter.

“V1” is a 2023 dense model. The paper does not spell this out, but the config file for DeepSeek’s November 2023 67B model has 95 layers, 8 key-value heads, a head dimension of 128 (8,192 hidden units over 64 heads) and bfloat16 weights. Multiply 95 layers by 2 (keys and values) by 8 heads by 128 by 2 bytes and you land on 389,120 exactly. That is my arithmetic, not a statement from the paper, but the match to the byte makes the identification safe. The baseline is a model with grouped-query attention from before the lab’s own cache-compression work began. It is a comparison of DeepSeek with its own earlier self, not with a contemporary rival.

It counts one memory term. The paper’s figure is the global cache, the part that sits in GPU memory and grows with context. It leaves out the sliding-window cache, which is bounded regardless of length. It also leaves out the separate persistent cache on SSD or host memory, which the paper reports as roughly one-eighth of V4-Flash’s. The conclusion says “at equal sequence lengths,” so this is a per-token size ratio and says nothing about a particular workload.

It stacks several generations of ideas. The four-fold step from V4-Flash to V4.1-Flash comes, per the paper, from cross-layer reuse of the cache and index plus FP4 storage for the main cache. The earlier steps combine latent-attention compression with sparse attention. The paper shows the step ratios but doesn’t apportion the 437 among techniques. The quality cost is also reported by the authors alone: “negligible” for one change, “marginal” for FP4. I found no independent reproduction of the 890-byte figure.

The blog also says DeepSeek’s earlier latent-attention work “compressed the cache by roughly 15x.” That matches the V2 paper’s headline of a 93.3% reduction against the 67B model (May 2024), which works out to about 15x by my arithmetic. The V4.1 chart puts V3.2 at only 8.1x below V1. I can’t tell from the documents how the two accountings differ, which is a reason to treat any single ratio as a statement under one counting convention. I made the same point about speed multiples: the denominator does half the work.

There is also a lineage that predates all of it. Noam Shazeer’s “One Write-Head is All You Need” (arXiv, 6 November 2019) proposed sharing keys and values across attention heads specifically to cut the memory traffic of incremental decoding. The chart DeepSeek draws is a descendant of that idea, four years before the 2023 baseline.

The proof

The blog’s second step: Claude Opus 5.5 cut its cache-read price by 60% against Opus 5, GPT-6.1 Sol by 80% against GPT-5.6 Sol’s late-July price, and since the caches got cheaper, the labs must have adopted DeepSeek’s design.

Both figures check out. Anthropic’s Opus 5.5 launch post (22 September) lists cache reads at $0.20 per million tokens, 60% below Opus 5’s $0.50. OpenAI’s pricing for GPT-6.1 Sol (29 September) lists cached input at $0.10 per million, against $0.50 for GPT-5.6 Sol when it launched in July.

Cache-read price cuts split into input price and discount multiplier Three dumbbells show cache-read price per million tokens before and after. Anthropic: Opus 5 at 50 cents to Opus 5.5 at 20 cents, from a 20 percent lower input price times a cache multiplier that halved from 10 to 5 percent of input. OpenAI against GPT-5.6 Sol in July: 50 cents to 10 cents, from a 60 percent lower input price times the same halving. OpenAI against GPT-6 Sol: 20 cents to 10 cents with the input price unchanged, so the halved multiplier accounts for the entire cut. CACHE-READ PRICE · USD PER MILLION TOKENS · HOLLOW = BEFORE, SOLID = AFTER $0 $0.30 $0.60 Anthropic OpenAI vs July OpenAI vs 22 Sep Opus 5 → Opus 5.5 GPT-5.6 Sol → 6.1 Sol GPT-6 Sol → 6.1 Sol $0.50 — Opus 5 cache read, Anthropic, 2026 $0.50 — GPT-5.6 Sol cache read, July 2026 pricing $0.20 — GPT-6 Sol cache read, 22 Sep 2026 $0.20 — Opus 5.5 cache read, Anthropic, 22 Sep 2026 $0.10 — GPT-6.1 Sol cache read, OpenAI, 29 Sep 2026 $0.10 — GPT-6.1 Sol cache read, OpenAI, 29 Sep 2026 $0.50 $0.20 $0.50 $0.10 $0.20 $0.10 input −20% × multiplier ×0.5 (10% → 5% of input) = −60% input −60% × multiplier ×0.5 = −80% input unchanged × multiplier ×0.5 = −50% Prices as listed by each lab; the decomposition is my arithmetic. GPT-5.6 Sol is the July launch price, before its 21 August cut.
In all three comparisons the cache discount, expressed as a share of the input price, went from 10% to 5%. In the third, nothing else moved.

The figure shows what the blog leaves out. The cache-read price is the input price times a discount multiplier, and at both labs the multiplier fell from 10% of input to 5%. Opus 5.5’s input price also dropped 20%, and 0.8 times 0.5 gives the 60% cut. GPT-6.1 Sol’s input price is $2 against $5 for GPT-5.6 Sol at launch, and 0.4 times 0.5 gives 80%. Against GPT-6 Sol, launched a week earlier, the input price is unchanged and the cache price simply halved. The two multiplier changes are visible on the price sheets without any architecture story.

That doesn’t show the labs didn’t change their serving. A halved multiplier is exactly what a cheaper cache would let a lab offer. It shows the price cut alone can’t tell the two stories apart, because a pricing decision produces the same sheet.

The primary sources don’t help the blog. I read both launch posts for any mention of KV caches, DeepSeek, or cache architecture and found none. Anthropic’s post says cache reads make up the majority of agentic and coding costs and describes Opus 5.5 as costing 40% less to run than Opus 5, with no mechanism given. OpenAI’s post frames its cut as room for agents that reuse context. The only cost explanation I found at OpenAI is for its 30 July cuts to GPT-5.6 Terra and Luna, which the announcement attributes to GPU kernel work and speculative decoding, six weeks before DeepSeek’s weights appeared. Its 21 August cut to GPT-5.6 Sol, a time-limited promotion, gives only a general line about improving efficiency.

Timeline of DeepSeek and Western releases, 21 August to 29 September 2026 A line from 21 August to 29 September with five marks. OpenAI cut GPT-5.6 Sol prices on 21 August. DeepSeek-V4.1-Flash's Hugging Face repository was created on 10 September. The DeepSeek paper was submitted to arXiv on 17 September. Anthropic released Opus 5.5 on 22 September, 12 days after the repository. OpenAI released GPT-6.1 Sol on 29 September, 19 days after. 2026 · SOLID = DEEPSEEK, HOLLOW = ANTHROPIC / OPENAI 10 Sep — DeepSeek-V4.1-Flash repository created on Hugging Face 17 Sep — DeepSeek paper, arXiv 2609.19969 v1 21 Aug — OpenAI cuts GPT-5.6 Sol prices 22 Sep — Claude Opus 5.5, Anthropic 29 Sep — GPT-6.1 Sol, OpenAI 21 Aug 10 Sep 17 Sep 22 Sep 29 Sep GPT-5.6 Sol price cut V4.1-Flash weights DeepSeek paper Opus 5.5 GPT-6.1 Sol Repository date is its creation date on Hugging Face; I could not confirm it was public that day.
Opus 5.5 arrived 12 days after the weights appeared and GPT-6.1 Sol 19 days after. That is possible if a lab had been working toward similar techniques, but not evidence of it.

Timing is the last argument the blog could make, and it is weak. Both Western releases came after DeepSeek published, by 12 and 19 days from the repository’s creation. A model’s serving stack and price are set on a longer schedule than that, and lower prices were already the direction of travel: OpenAI’s July and August cuts predate DeepSeek’s release entirely.

What survives

The claim that holds up is a DeepSeek claim: by its own count, its newest model’s global cache is 437 times smaller per token than its 2023 dense model’s, and four times smaller than its April model’s. That is a real engineering result, and the paper is public. It tells you about DeepSeek’s trajectory and says little about anyone else’s.

The adoption claim needs something the price sheets can’t supply: a statement from a lab about its own architecture, or a measurement from outside. Neither lab’s launch post offers either, and I found no independent audit of DeepSeek’s cache figure either. A halved discount multiplier at two vendors within a week of each other is a fact. It doesn’t tell us what’s inside the models.

References

  1. DeepSeek-AI (2026). DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression. arXiv:2609.19969v1, 17 September 2026.
  2. insufferable.dev (2026; no named author). The AI Race Just Got Awkward. insufferable.dev, 29 September 2026.
  3. Anthropic (2026). Introducing Claude Opus 5.5. 22 September 2026.
  4. OpenAI (2026). Introducing GPT-6.1 Sol. 29 September 2026.
  5. OpenAI (2026). GPT-5.6. 9 July 2026; update note of 21 August 2026.
  6. OpenAI (2026). Announcing a major Price drop for 5.6 Terra and Luna and Fast mode for 5.6-Sol. OpenAI Developer Community, 30 July 2026.
  7. OpenAI (2026). API changelog. Entry of 29 September 2026.
  8. Hugging Face (2026). deepseek-ai/DeepSeek-V4.1-Flash. Repository created 10 September 2026 per the Hugging Face API.
  9. DeepSeek-AI (2023). deepseek-llm-67b-base config.json. Hugging Face; the “V1” baseline, used for the 389,120-byte derivation.
  10. DeepSeek-AI (2024). DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434, 7 May 2024.
  11. Shazeer, N. (2019). Fast Transformer Decoding: One Write-Head is All You Need. arXiv:1911.02150, 6 November 2019.
  12. Gautam Parab (2026). Jev Is 5 to 200 Times Faster, Depending on the Denominator. 20 September 2026.