A Lot More B Than A

A story titled “From the creator of Redis; run LLM locally with ds4” sat on the Hacker News front page this afternoon, close to 90 points within a few hours. The link goes to dwarfstar.sh, a clean, well-built site with a hardware-fit calculator and a benchmark table. It is not, in fact, from the creator of Redis. Its own About page says so: “A community site for DwarfStar 4… Built as a community-maintained web surface around ds4.” Salvatore Sanfilippo (antirez, who did write the thing) has his own blog and a GitHub repository, neither of which is what today’s story links to.

That’s a small slip, and an unintentional one; the community site is good, it credits antirez correctly throughout, and it says plainly that it is not the source. But it’s a fitting start to a day of looking at this project, because the more interesting question underneath it is the same shape: who’s actually making the claim, and does the thing being cited for it say what it’s cited for.

The claim, and how hedged it actually is

DwarfStar 4 (ds4) is a local inference engine, mostly C, that antirez built to run DeepSeek V4 Flash, and since then GLM 5.x and Qwen3.8 Flash Next, on a single high-memory Mac, an Nvidia DGX Spark, or an AMD Strix Halo box. It isn’t new; he shipped it in May and wrote about it on his own site in a post called “A few words on DS4,” which went to the Hacker News front page itself that month with 440 points. Today’s story is a second, smaller wave, pointed at the documentation rather than the thing.

The sentence everyone quotes from the May post is this one, in full:

“It is the first time since I play with local inference (I play with it since the start) that I find myself using a local model for serious stuff that I would normally ask to Claude / GPT. This, I think, is really a big thing.”

Read closely, it’s a narrower claim than the summary version that circulates. “The first time,” not “now routinely.” “Serious stuff,” left undefined, not a named task or workload. “I think” is his own hedge, not a measurement. He follows it with a comparison, not a verdict: “If you can imagine in your mind the small good local model experience as A, and the frontier model you use online as B, DS4 is a lot more B than A.” Closer to B. Not B.

What the quantization actually does

The mechanism behind that closeness is described on dwarfstar.sh in almost these words: “Asymmetric quantization targets the routed experts while preserving critical paths.” The GitHub README itself is plainer: “DeepSeek V4 Flash and PRO, GLM 5.2, tolerate aggressive routed-expert quantization.” DeepSeek V4 Flash is a large mixture-of-experts model: 284 billion parameters, per dwarfstar.sh, though Hugging Face’s own metadata tag for the base checkpoint says 292B, and neither page explains the gap. Most of a MoE model’s weight sits in its routed experts, the pieces that get swapped in per token; ds4 compresses those hard and leaves the shared and attention paths at higher precision. Antirez’s own phrase for the ratio, in the May post, is “an extremely asymmetric quants recipe of 2/8 bit.” It’s the same lab whose memory-efficiency claims for its hosted models needed a similar fine-print check when they made the rounds last month.

The asymmetric quantization recipe behind ds4 Horizontal bar chart. Shared and attention paths are kept at 8-bit precision, a bar four times as long as the one for routed experts, which hold most of a mixture-of-experts model's parameters and are compressed to 2-bit. DS4 QUANTIZATION RECIPE · PER THE GITHUB README AND ANTIREZ'S OWN POST Shared / attention paths Routed experts (bulk of the parameters) 8-bit — shared and attention paths, kept precise. ds4 README. 2-bit — routed experts, compressed. ds4 README; antirez: "2/8 bit". 8-bit 2-bit Bar length scaled to bits per weight. "Compressed, not lobotomized" is dwarfstar.sh's own phrase for this trade.
The asymmetry is the whole engineering bet: most of the model's weight goes to 2 bits, the parts that most affect output stay at 8.

The hardware floor needs a small correction too. The project’s own landing page advertises “64 GB+ depending on model,” which is true for the smaller Qwen3.8 build but not for the model the project is named around: the README is explicit that DeepSeek V4 Flash itself wants “Macs with 96 GB or more,” with smaller machines falling back to SSD streaming. Antirez’s own figure from May is “96 or 128GB,” two specific tiers, not the 64GB floor the current landing page leads with.

ds4 is not a small side project at this point. The repository carries 22.9k GitHub stars, 797 commits, 55 listed contributors, 279 open issues, 474 open pull requests, and, deliberately, no tagged release. The About page on the community site explains why: “there are deliberately no GitHub releases or tags, and docs are written to be useful to agents as much as to humans.” The README calls the software “currently very fast changing… beta quality.”

What the benchmark table shows, read as two machines rather than one number

The project’s own speed table, reproduced on dwarfstar.sh from the repository’s benchmark data, reads a bit differently depending on which axis you hold fixed.

ds4 throughput, M5 Max versus DGX Spark, at 2,048 and 65,536-token context Two small-multiple slope charts, both at q2 quantization on 128GB machines. Prefill: M5 Max falls from 790.2 to 398.5 tokens per second as context grows from 2,048 to 65,536 tokens, while DGX Spark barely moves, from 825.8 to 823.0. Generation: M5 Max falls from 39.4 to 27.6 tokens per second, DGX Spark from 18.1 to 13.8. DS4 BENCHMARK TABLE · Q2 QUANTIZATION · BOTH MACHINES 128GB Prefill (tokens/sec) Generation (tokens/sec) 790.2 t/s — M5 Max 128GB, prefill, 2,048-token context 398.5 t/s — M5 Max 128GB, prefill, 65,536-token context 39.4 t/s — M5 Max 128GB, generation, 2,048-token context 27.6 t/s — M5 Max 128GB, generation, 65,536-token context 825.8 t/s — DGX Spark 128GB, prefill, 2,048-token context 823.0 t/s — DGX Spark 128GB, prefill, 65,536-token context 18.1 t/s — DGX Spark 128GB, generation, 2,048-token context 13.8 t/s — DGX Spark 128GB, generation, 65,536-token context 790.2 398.5 39.4 27.6 825.8 823.0 18.1 13.8 2,048 tok 65,536 tok 2,048 tok 65,536 tok M5 Max 128GB (solid) DGX Spark 128GB (dashed) Figures from dwarfstar.sh's reference table, both machines at q2. Not an independent benchmark.
Read by model, the headline is a near-halving at long context. Read by machine, the Mac does almost all of that halving; the Spark's prefill speed barely moves.

Flatten the table to a single number and the story is “generation speed drops as context grows,” which is unsurprising and true of basically every model on every machine. Split it by hardware and a sharper fact appears: the DGX Spark’s prefill speed is nearly flat across a 32x jump in context length (825.8 down to 823.0 tokens per second), while the M5 Max’s prefill speed falls by nearly half, a 49.6% drop, 790.2 to 398.5. Generation speed drops on both machines by a roughly comparable amount, about 24% on the Spark and about 30% on the Mac, much closer to each other than the prefill gap would predict. These are the project’s own numbers, not an outside audit, and the table doesn’t say why the two machines diverge on prefill the way they do.

The argument happening in the comments, checked

The liveliest thread under today’s post is a dispute about whether the 2-bit quantization is actually any good. User doctorpangloss: “the problem is the dsv4 checkpoint so quantized isn’t very good.” User c0rruptbytes, in reply: “not my experience — the ds4 quants were very good beating the unsloth quants,” linking to a benchmark report on GitHub. User dotancohen, unconvinced: “That’s quite the statement — unsloth quants are amazing.”

Screenshot of a Hacker News comment thread. doctorpangloss writes 'the problem is the dsv4 checkpoint so quantized isn't very good.' ilaksh asks which checkpoint was tested. c0rruptbytes replies 'not my experience, the ds4 quants were very good beating the unsloth quants' with a link to a GitHub benchmark. dotancohen replies 'That's quite the statement - unsloth quants are amazing.'
The exchange as it stood a few hours after the story posted. Nobody in the visible thread settles it with the checkpoint and quant level ilaksh asked for. Image: Hacker News, comment thread on "From the creator of Redis; run LLM locally with ds4," news.ycombinator.com/item?id=49936575. Reproduced for commentary.

I read the linked report. It’s by Michael Asper, titled “Running DeepSeek V4 Flash 0731 locally on SlopCodeBench,” and it does not mention ds4, DwarfStar, or Unsloth anywhere in its text; I checked the raw file for all three strings and got nothing. What it actually reports is two runs of the same quantized model a day apart, under the pi coding agent, where correctness jumped from 1 of 17 to 5 of 17 solved checkpoints. The author’s own framing, in the opening paragraph: “the two reports are not a quantization A/B; too much differs.” Two things changed between them at once, the output-token cap (32,768 to 49,152) and the served quant (a vendor “UD-Q2_K_XL” build to “a larger community q2q4 imatrix quant”), and the piece’s whole point is that you can’t credit either change alone. Its own conclusion: “Whether the new quant also writes better code cannot be answered from these two runs.” That sentence is the direct answer to the claim it was cited to support, and the answer is no, not from this data.

This doesn’t read like a bad-faith citation. It reads like someone skimming a benchmark post for its headline correctness jump and remembering it as “the new quant won,” which isn’t a wild misreading of a 1-to-5 jump. It’s just not what the report, read past its table, says it measured. The doctorpangloss claim is just as unresolved: nobody in the visible thread names a specific checkpoint or quant level, which is exactly what ilaksh asked for. As of this writing, the question that opened the thread, is ds4’s quantization good, is still open, and the one piece of “proof” offered for either side doesn’t hold up as proof of either.

Built fast, with help

There’s a second tension here, and antirez doesn’t hide it. From the May post: the local-AI movement’s accumulated know-how “can be leveraged more promptly because of GPT 5.5 (otherwise you can’t build DS4 in one week, and even with all this help you need to know how to gently talk to LLMs).” He worked 14-hour days that week, against what he describes as his normal 4-to-6. The current README carries its own section, “AI full disclosure”: “This software is developed with strong assistance from AI coding agents and with humans leading the ideas, testing, and debugging. We say this openly because it shaped how the project was built.” The same README is candid that ds4 isn’t built from nothing either: “ds4.c does not link against GGML, but it exists thanks to the path opened by the llama.cpp project,” and it names specific pieces kept under llama.cpp/GGML’s MIT license, “GGUF quant layouts and tables, CPU quant/dot logic, and certain kernels.” A tool built in a week, with a hosted frontier model’s help, on top of an acknowledged debt to an existing open-source runtime, whose purpose is to reduce how often you need a hosted frontier model. None of that is a secret (antirez states all three facts himself), but it’s a more complicated picture than “narrow C engine, built from scratch, replaces the cloud,” which is roughly how it reads on the fastest pass through the landing page.

It also isn’t out of character. In March 2007, antirez released Picol, a Tcl-alike interpreter he describes in its own README as “500 lines of code,” adding that he revisited and rewrote it as recently as February of this year, a few months before DS4. The instinct is old, close to twenty years: pick one target, write a small, legible engine from scratch rather than wrap something general, make the source worth reading. DS4 is the same bet at a much larger scale, aimed at a 284-billion-parameter model instead of a toy language.

What’s actually shown

The parts of this story that hold up on inspection: a working, fast-moving open-source project with tens of thousands of stars and no tagged release; a documented, specific quantization recipe that the project’s own maintainers describe plainly, including its tradeoffs; a benchmark table that, read by machine rather than by headline, shows an uneven and only partly explained hardware story. The parts that don’t hold up to the weight put on them today: the Hacker News title’s attribution, the community site’s implicit claim to be the primary source, and, most concretely, a benchmark cited mid-argument as proof of a quantization-quality verdict it explicitly declines to reach. What’s left, once those are set aside, is smaller and more honest than either side of today’s thread: one engineer, in May, saying this was the first time, for him, that a local model did real work he’d otherwise have sent to Claude or GPT. Not a claim about ds4 in general, and not the number that got cited to settle an argument it never addressed, which is a pattern this series keeps running into: a figure that’s accurate on its own terms, attached after the fact to a claim it was never measuring. A claim about one week, for one person, who also happens to be the person who wrote it.

References

  1. dwarfstar.sh (community site, 2026). DwarfStar 4 (ds4): Local DeepSeek V4.1, Qwen and GLM. Built 17 September 2026 per the site’s own project card.
  2. dwarfstar.sh (community site, 2026). About DwarfStar 4.
  3. Sanfilippo, S. (“antirez”) (2026). A few words on DS4. antirez.com, approximately 14–15 May 2026 (dated by the page’s relative-time counter and a matching Hacker News submission timestamp).
  4. Sanfilippo, S. and contributors (2026). antirez/ds4, README, “DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm.” GitHub, accessed 2 October 2026.
  5. fibo (submitter) (2026). From the creator of Redis; run LLM locally with ds4. Hacker News, 2 October 2026.
  6. Asper, M. (2026). Running DeepSeek V4 Flash 0731 locally on SlopCodeBench. GitHub, runs dated 7–8 August 2026.
  7. Hugging Face (2026). deepseek-ai/DeepSeek-V4-Flash-Base. Model metadata tag: 292B params; no model card published.
  8. Sanfilippo, S. (2026). antirez/picol, README, “A Tcl interpreter in 500 lines of code,” originally released 15 March 2007, revised February 2026.
  9. Gautam Parab (2026). 437x Against a 2023 Baseline: What the DeepSeek Cache Number Does and Doesn’t Show. 1 October 2026.
  10. Gautam Parab (2026). Jev Is 5 to 200 Times Faster, Depending on the Denominator. 20 September 2026.