← Gautam Parab

Two of the Seven Judges Were Claude Code

On 7 September the Latent Space Frontier AEO tracker published what seven frontier AI agents recommend when you ask them to pick a product. 161 categories, six phrasings of each buyer scenario, 6,762 saved answers, every prompt and answer pair inspectable. In the coding agents category the winner is Claude Code, with a 57% net score and 14 first choices out of 42 answers. Second is OpenAI Codex at 44% and 12 first choices. Third is Cursor at 38% and 6.

The write-up notices the obvious thing and shrugs at it: “when models are asked for coding agent recommendations, Fable/Opus like Claude Code and Sol/Astra like Codex and Grok loves Cursor and Muse loves Muse Code and SWE-1.7 loves Devin and so on. I wonder why.”

I want to take the wondering seriously, because the answer is written down on the tracker’s own methodology page, and it is not the one that sentence implies.

The seven things surveyed are not seven models. They are seven model-CLI configurations: “Sol and Astra through Codex; Fable and Opus through Claude Code; Grok through Cursor; Spark through Muse Code; SWE through Devin.” Two of the seven judges in the coding agents category were Claude Code. Two were Codex. One was Cursor, one was Muse Code, one was Devin. Five of the six products at the top of that leaderboard held at least one seat on the panel, and the two holding two seats finished first and second.

Do the arithmetic slowly. Six frames times seven configurations is 42 answers, six per configuration. So Anthropic’s two configurations could account for at most 12 of Claude Code’s 14 first choices, and OpenAI’s two for at most 12 of Codex’s 12. Cursor took exactly 6 first choices and had exactly one configuration. Muse Code took exactly 6 and had exactly one. Devin took 4 of a possible 6.

First choices observed against the ceiling each product's own panel seats could supply A dumbbell chart of six coding-agent products in the Latent Space AEO tracker's coding agents category, September 2026. Each product has a tick marking the maximum first choices its own vendor's configurations could have produced, at six answers per seat, and a filled dot marking the first choices actually observed out of 42 answers. Claude Code held two seats, a ceiling of 12, and was observed with 14 first choices, above its own ceiling. OpenAI Codex held two seats, ceiling 12, observed 12, exactly at the ceiling. Cursor held one seat, ceiling 6, observed 6, at the ceiling. Muse Code held one seat, ceiling 6, observed 6, at the ceiling. Devin held one seat, ceiling 6, observed 4, below the ceiling. GitHub Copilot held no seat, ceiling 0, observed 1. CODING AGENTS · FIRST CHOICES OF 42 ANSWERS · AEO TRACKER, SEPT 2026 Claude Code OpenAI Codex Cursor Muse Code Devin GitHub Copilot 2 seats 2 seats 1 seat 1 seat 1 seat no seat Ceiling 12 — two own-harness seats, six answers each Ceiling 12 — two own-harness seats, six answers each Ceiling 6 — one own-harness seat Ceiling 6 — one own-harness seat Ceiling 6 — one own-harness seat Ceiling 0 — no own-harness seat 14 first choices of 42 — Claude Code 12 first choices of 42 — OpenAI Codex 6 first choices of 42 — Cursor 6 first choices of 42 — Muse Code 4 first choices of 42 — Devin 1 first choice of 42 — GitHub Copilot 14 12 6 6 4 1 0 4 8 12 16 Tick = ceiling its own configurations could supply (6 answers per seat). Dot = first choices observed. Claude Code is the only seat-holder above its own ceiling, so at least 2 of its 14 came from elsewhere.
Four of the six ranked products have a first-choice tally that fits inside what their own configuration could have produced. Claude Code does not, and 14 is larger than 12.

None of this proves each configuration voted for its own house. It supports the weaker claim: of the products that held seats, only Claude Code demonstrably won votes from outside its own harness, because 14 is larger than 12.

The tracker is careful about exactly this. Its methodology page states it directly: “Native CLI harnesses, with search available where supported. Retrieval and harness context are part of the measured configuration.” And under a heading that reads “Model comparisons include the harness”: “Tools, context and retrieval differ between configurations. A Sol-Astra or Opus-Fable difference describes the recorded setup; it does not isolate the effect of model weights.”

It goes further, and discloses something most studies would bury: “an audit found personal skill access in some Cursor runs and potential account context in Claude.” A model asked which coding agent to use, while running inside a coding agent, carrying that agent’s system prompt and tool definitions and in some runs a personal account context, is not a neutral witness. The study knows this and says so. The write-up says “I wonder why.”

That matters more than a methodological quibble, because the 2026 literature has converged on something specific about when self-preference appears. The most careful study I found this year separates the label from the content, evaluating choices that “carry no stylistic fingerprint” so a judge cannot infer authorship from the text. Under blind evaluation the effect disappears, with what the authors describe as a small remainder in the opposite direction. Labelling an identical-quality answer as the judge’s own, without naming any model, still shifts scores in both directions (Chae et al., arXiv:2608.18091, June 2026, EMNLP Findings 2026). On that evidence self-preference is largely a response to a salient authorship cue rather than a buried taste for one’s own prose.

Running a model inside its vendor’s own CLI is that cue at maximum, applied to the one category where it names a product on the ballot.

Nor does frontier capability wash it out. A study covering 20 models built equal-quality response pairs specifically to separate a judge’s discriminating ability from its bias, and reported that “advanced capabilities are often uncorrelated, or even negatively correlated, with low SPB” (Yang et al., arXiv:2604.22891, April 2026). The better model is not the more disinterested one.

The number that travelled wrong

There is a second gap between the tracker and its own coverage, and this one is arithmetic rather than framing.

The write-up: “There are 28 categories (out of our total 161) which have a universally dominant primary choice - among all surveyed frontier models.”

The tracker’s findings page, describing the same 28: “That is 28 of 61 categories with complete coverage and one clear leader per model.” Immediately below the agreement table: “100 categories are excluded from this agreement test because a model has incomplete coverage, no named first choice, or tied leaders.”

Why the same 28 categories support two very different agreement rates A three-stage funnel from the Latent Space AEO tracker, September 2026. All 161 categories narrow to 61 categories eligible for the agreement test, after 100 are excluded for incomplete coverage, no named first choice, or tied leaders. Of those 61 eligible categories, 28 have the same leader across all seven configurations. Dividing 28 by the full 161, the denominator used in the write-up, gives 17 percent. Dividing 28 by the 61 eligible categories, the denominator on the tracker's own findings page, gives 46 percent. AGREEMENT TEST · DENOMINATORS · AEO TRACKER, SEPT 2026 All categories Eligible for the test Same leader, all seven 161 categories surveyed — AEO tracker, September 2026 61 categories with complete coverage and one clear leader per model 28 categories with the same leader across all seven configurations 161 61 28 100 excluded: incomplete coverage, no named first choice, or tied leaders 28 of 161 17% 28 of 61 46% the denominator used in the write-up the denominator on the tracker's findings page Same 28 categories, two denominators. The version that travelled is the one that makes agreement look rare.
The excluded 100 are not disagreements. They are categories where the test could not run.

28 of 161 is 17%. 28 of 61 is 46%. The version that travelled understates the real agreement rate by a factor of about 2.6, and the correction cuts against the sceptical reading, which is usually a sign a correction is sound. Where all seven configurations produced a clean answer, they landed on the same leader almost half the time. pnpm, Sentry, PostgreSQL, Vitest, Slack, Playwright, Instructor and Ramp are unanimous, each at a 100% first-choice rate across every configuration.

What the design can isolate

Strip the harness confound out of the headline and something better is left, because the design contains its own control.

Sol and Astra are two OpenAI models running through the same Codex harness. Fable and Opus are two Anthropic models running through the same Claude Code harness. Within each pair the tools, the system prompt and the retrieval stack are held fixed and only the weights change. That is the comparison the study supports, and the tracker leads with it: Sol and Astra “have different leaders in 33 of 121 comparable categories.” In AI sandboxes the flip is total. Sol picks Modal in all six frames, Astra picks E2B in all six.

Roughly a quarter of categories change hands when nothing moves but the model generation inside the same product. So the harness does not explain everything. The study simply cannot separate harness from weights in its headline claim, and where it can separate them, the weights alone move about a quarter of the answers.

Six phrasings are not six trials

The tracker is equally clear about a limit that the coverage turns into a selling point.

“Six phrasings are six observations. They retain the same buyer scenario. With one response per model/frame, wording, retrieval, timing and randomness can all affect a difference. This is not a repeated-trial estimate of a stable preference.”

The write-up’s version: “Astra is FAR less likely to change its mind when you lightly paraphrase your question. This makes the value of AEO itself rise as choice randomness declines.”

The underlying number is interesting. Astra repeats its leader across all six frames in 54 categories, against Sol’s 34. But with one sample per frame and no same-condition repeat baseline, frame-to-frame variation and run-to-run variation are the same measurement. A model that is simply more deterministic in effect will look more decisive on this metric without holding a more stable preference underneath. And stability does not track model recency in the way the inference assumes: a 2026 benchmark built for exactly this question reports that “the largest and newest models are not the most consistent” under paraphrase (Bellibatlu et al., arXiv:2604.23478, April 2026).

You cannot get from these data to the claim that a vendor’s optimisation work will move the ranking. The tracker says so directly: the study “cannot establish market share, product quality, causal effects of wording, or whether a vendor action will change a recommendation.” That last clause is the entire premise of answer engine optimisation, disclaimed by the study being sold as evidence for it.

One external check is available, and it is unflattering to the idea that these rankings track the market.

JetBrains surveyed more than 15,000 professional developers between May and July 2026 and reported Claude Code at 39% adoption globally and 47% in the US, GitHub Copilot at 21% and falling from 29% year over year, OpenAI Codex at 16% and rising from 3%, and Cursor at 12%, down from 18% (JetBrains, August 2026). JetBrains sells a competing product, so treat the framing as interested; the sample is large and the fielding dates are stated.

Recommendation rate against developer adoption for four coding agents A scatter plot of four coding agents. The horizontal axis is the share of professional developers reporting use of the tool at work, from the JetBrains survey of more than 15,000 developers fielded May to July 2026. The vertical axis is the share of the AEO tracker's 42 coding-agent answers in which the tool was the first choice, September 2026. Claude Code, holding two panel seats, sits at 39 percent adoption and a 33 percent first-choice rate. OpenAI Codex, two seats, sits at 16 percent adoption and 29 percent. Cursor, one seat, sits at 12 percent adoption and 14 percent. GitHub Copilot, which held no seat on the panel, sits at 21 percent adoption, the second highest of the four, and a 2 percent first-choice rate, the lowest. The two axes measure different things and the comparison is of ordering, not of levels. RECOMMENDED VS USED · AEO TRACKER SEPT 2026 · JETBRAINS MAY–JUL 2026 Claude Code — 39% adoption, 33% first-choice rate, 2 panel seats OpenAI Codex — 16% adoption, 29% first-choice rate, 2 panel seats Cursor — 12% adoption, 14% first-choice rate, 1 panel seat GitHub Copilot — 21% adoption, 2% first-choice rate, no panel seat Claude Code OpenAI Codex Cursor GitHub Copilot 2 seats 2 seats 1 seat no seat 0% 10% 20% 30% 40% 0% 10% 20% 30% Vertical: share of the 42 answers where the tool was first choice Horizontal: share of surveyed professional developers using the tool at work Two different measures on two different populations; read the ordering, not the levels.
The second most widely used agent in the survey held no seat on the panel and took one first choice out of 42.

Copilot is the awkward case. By adoption it is second of the four. On the panel it had no seat, and it took a single first choice out of 42, behind every product that did have one. I would not lean hard on this: the two axes measure different things on different populations, the tracker flags name ambiguity on Copilot’s entry, and low adoption momentum is a real alternative explanation. But if you were looking for a pattern where recommendation rank tracks who was in the room rather than who is being used, this is what it would look like.

None of this is new as a phenomenon, only as a substrate. In 1994 a group at the Brockton/West Roxbury VA Medical Center pulled every randomised trial of non-steroidal anti-inflammatories in arthritis indexed by MEDLINE between September 1987 and May 1990, narrowed to 61 qualifying articles, and had reviewers blinded to sponsorship read the narrative conclusions. 56 of the trials were manufacturer-associated. The manufacturer’s own drug was reported as comparable with the comparison drug in 71.4% of them and superior in 28.6%, which is to say it came out equal or better in all 56. The efficacy claims were usually backed by the trial data. The safety claims were not: among the 22 trials naming one drug as less toxic, the sponsor’s drug was the safer one 86.4% of the time, and the narrative interpretation was justified by the data in 12 of those 22 (Rochon et al., Arch Intern Med 154(2):157-163, 1994). Fifty-six for fifty-six is a cleaner result than anything in the AEO tracker, and it took the field two decades of registries and disclosure rules to do much about it.

The regulatory vocabulary for the structural version of this already exists. In June 2017 the European Commission fined Google 2.42 billion euros after finding, in the Commissioner’s words, that it “abused its market dominance as a search engine by promoting its own comparison shopping service in its search results, and demoting those of competitors” (European Commission, IP/17/1784). That was a ranked list of links. A coding agent recommending a coding agent is a ranked list of one, delivered by a product with a commercial interest in the answer, with no disclosure surface and no ranking to inspect.

What would settle it

The fix is not complicated and it is cheap next to what this study already cost. Query each model twice: once through its vendor’s own CLI, once through a neutral harness with a common system prompt and the same retrieval stack. The difference between those two numbers is the harness effect, and it is the number everyone wants. Then repeat each frame several times per condition so that frame variance can be told apart from run variance. The tracker has already built the paired-comparison machinery for something else, an earlier search-on against search-off experiment of 80 answers across 16 scenarios and five configurations, so the method is in hand.

The expensive part is done. Every prompt and answer pair is inspectable, which is more than almost any evaluation of this kind offers, and the methodology page volunteers limitations that would have been easy to omit. The problem is not the study. It is that the study’s own caveats did not survive the trip into the write-up, and the write-up is what gets cited.

One more entry, offered without a conclusion attached: among the tracker’s 161 editorially selected categories, the AI Podcasts category is led by Latent Space at 85%, and the AI Newsletters category is led by Latent Space at 52%.

I have written before about a pair of preprints whose viral summary said close to the opposite of what they said, and about benchmark margins that shrink when you read the primary table instead of the announcement. This is the third instance of the same shape in three days. Each time the primary source was more careful than its summary. Each time the summary is what moved.

References

  1. Latent Space. Frontier AEO Tracker. Snapshot labelled September 2026. All tracker figures here are read from its own findings, methodology and category pages rather than from coverage.
  2. Latent Space. (2026, September 7). The Frontier AEO Tracker: What Astra Chooses.
  3. JetBrains. (2026, August). State of AI Coding Agents. A vendor-run survey.
  4. Chae et al. (2026, June). arXiv:2608.18091. EMNLP Findings 2026. Read from the arXiv record itself.
  5. Yang et al. (2026, April). arXiv:2604.22891. Read from the arXiv record itself.
  6. Bellibatlu et al. (2026, April). arXiv:2604.23478. Read from the arXiv record itself.
  7. Rochon et al. (1994). Arch Intern Med, 154(2), 157–163. Cited and dated as historical context; read at the primary source.
  8. European Commission. (2017). Press release IP/17/1784. Cited and dated as historical context; read at the primary source.