A post on insufferable.dev, dated 29 September, argues that the Western labs are now adopting Chinese labs’ work rather than distilling it. Its evidence has two parts: a 437x drop in KV-cache size from DeepSeek, and two cache-read price cuts at Anthropic and OpenAI that it calls “the proof of the adoption.” The post has no personal byline, and it links to nothing that backs these figures: not the paper, not either price page.
I checked both parts. The first is accurate and narrower than it sounds. The second is arithmetically correct and doesn’t prove what it is offered to prove.
Where 437 comes from
The number is in DeepSeek’s own paper on V4.1-Flash (arXiv 2609.19969, submitted 17 September). Figure 1(b) charts “global KV cache per token” in bytes for four models: V1 at 389,120, V3.2 at 48,068, V4-Flash at 3,514, and V4.1-Flash at 890. The caption states the reductions against V4-Flash and V1 as roughly 4-fold and 437-fold. Divide 389,120 by 890 and you get 437.2. The blog’s arithmetic is fine.
Three things about the baseline matter.
“V1” is a 2023 dense model. The paper does not spell this out, but the config file for DeepSeek’s November 2023 67B model has 95 layers, 8 key-value heads, a head dimension of 128 (8,192 hidden units over 64 heads) and bfloat16 weights. Multiply 95 layers by 2 (keys and values) by 8 heads by 128 by 2 bytes and you land on 389,120 exactly. That is my arithmetic, not a statement from the paper, but the match to the byte makes the identification safe. The baseline is a model with grouped-query attention from before the lab’s own cache-compression work began. It is a comparison of DeepSeek with its own earlier self, not with a contemporary rival.
It counts one memory term. The paper’s figure is the global cache, the part that sits in GPU memory and grows with context. It leaves out the sliding-window cache, which is bounded regardless of length. It also leaves out the separate persistent cache on SSD or host memory, which the paper reports as roughly one-eighth of V4-Flash’s. The conclusion says “at equal sequence lengths,” so this is a per-token size ratio and says nothing about a particular workload.
It stacks several generations of ideas. The four-fold step from V4-Flash to V4.1-Flash comes, per the paper, from cross-layer reuse of the cache and index plus FP4 storage for the main cache. The earlier steps combine latent-attention compression with sparse attention. The paper shows the step ratios but doesn’t apportion the 437 among techniques. The quality cost is also reported by the authors alone: “negligible” for one change, “marginal” for FP4. I found no independent reproduction of the 890-byte figure.
The blog also says DeepSeek’s earlier latent-attention work “compressed the cache by roughly 15x.” That matches the V2 paper’s headline of a 93.3% reduction against the 67B model (May 2024), which works out to about 15x by my arithmetic. The V4.1 chart puts V3.2 at only 8.1x below V1. I can’t tell from the documents how the two accountings differ, which is a reason to treat any single ratio as a statement under one counting convention. I made the same point about speed multiples: the denominator does half the work.
There is also a lineage that predates all of it. Noam Shazeer’s “One Write-Head is All You Need” (arXiv, 6 November 2019) proposed sharing keys and values across attention heads specifically to cut the memory traffic of incremental decoding. The chart DeepSeek draws is a descendant of that idea, four years before the 2023 baseline.
The proof
The blog’s second step: Claude Opus 5.5 cut its cache-read price by 60% against Opus 5, GPT-6.1 Sol by 80% against GPT-5.6 Sol’s late-July price, and since the caches got cheaper, the labs must have adopted DeepSeek’s design.
Both figures check out. Anthropic’s Opus 5.5 launch post (22 September) lists cache reads at $0.20 per million tokens, 60% below Opus 5’s $0.50. OpenAI’s pricing for GPT-6.1 Sol (29 September) lists cached input at $0.10 per million, against $0.50 for GPT-5.6 Sol when it launched in July.
The figure shows what the blog leaves out. The cache-read price is the input price times a discount multiplier, and at both labs the multiplier fell from 10% of input to 5%. Opus 5.5’s input price also dropped 20%, and 0.8 times 0.5 gives the 60% cut. GPT-6.1 Sol’s input price is $2 against $5 for GPT-5.6 Sol at launch, and 0.4 times 0.5 gives 80%. Against GPT-6 Sol, launched a week earlier, the input price is unchanged and the cache price simply halved. The two multiplier changes are visible on the price sheets without any architecture story.
That doesn’t show the labs didn’t change their serving. A halved multiplier is exactly what a cheaper cache would let a lab offer. It shows the price cut alone can’t tell the two stories apart, because a pricing decision produces the same sheet.
The primary sources don’t help the blog. I read both launch posts for any mention of KV caches, DeepSeek, or cache architecture and found none. Anthropic’s post says cache reads make up the majority of agentic and coding costs and describes Opus 5.5 as costing 40% less to run than Opus 5, with no mechanism given. OpenAI’s post frames its cut as room for agents that reuse context. The only cost explanation I found at OpenAI is for its 30 July cuts to GPT-5.6 Terra and Luna, which the announcement attributes to GPU kernel work and speculative decoding, six weeks before DeepSeek’s weights appeared. Its 21 August cut to GPT-5.6 Sol, a time-limited promotion, gives only a general line about improving efficiency.
Timing is the last argument the blog could make, and it is weak. Both Western releases came after DeepSeek published, by 12 and 19 days from the repository’s creation. A model’s serving stack and price are set on a longer schedule than that, and lower prices were already the direction of travel: OpenAI’s July and August cuts predate DeepSeek’s release entirely.
What survives
The claim that holds up is a DeepSeek claim: by its own count, its newest model’s global cache is 437 times smaller per token than its 2023 dense model’s, and four times smaller than its April model’s. That is a real engineering result, and the paper is public. It tells you about DeepSeek’s trajectory and says little about anyone else’s.
The adoption claim needs something the price sheets can’t supply: a statement from a lab about its own architecture, or a measurement from outside. Neither lab’s launch post offers either, and I found no independent audit of DeepSeek’s cache figure either. A halved discount multiplier at two vendors within a week of each other is a fact. It doesn’t tell us what’s inside the models.
References
- DeepSeek-AI (2026). DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression. arXiv:2609.19969v1, 17 September 2026.
- insufferable.dev (2026; no named author). The AI Race Just Got Awkward. insufferable.dev, 29 September 2026.
- Anthropic (2026). Introducing Claude Opus 5.5. 22 September 2026.
- OpenAI (2026). Introducing GPT-6.1 Sol. 29 September 2026.
- OpenAI (2026). GPT-5.6. 9 July 2026; update note of 21 August 2026.
- OpenAI (2026). Announcing a major Price drop for 5.6 Terra and Luna and Fast mode for 5.6-Sol. OpenAI Developer Community, 30 July 2026.
- OpenAI (2026). API changelog. Entry of 29 September 2026.
- Hugging Face (2026). deepseek-ai/DeepSeek-V4.1-Flash. Repository created 10 September 2026 per the Hugging Face API.
- DeepSeek-AI (2023). deepseek-llm-67b-base config.json. Hugging Face; the “V1” baseline, used for the 389,120-byte derivation.
- DeepSeek-AI (2024). DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434, 7 May 2024.
- Shazeer, N. (2019). Fast Transformer Decoding: One Write-Head is All You Need. arXiv:1911.02150, 6 November 2019.
- Gautam Parab (2026). Jev Is 5 to 200 Times Faster, Depending on the Denominator. 20 September 2026.