Gemini 4 Argon Is Held Back Over Cyber, and Its Results Table Has One Cyber Row

Google announced Gemini 4 Argon on 30 September, and in the same post said it is “rolling out to a set of trusted cyber defenders through our Fairwind Program.” For those defenders, and for Google’s own teams, the post says the model ships “without cyber guardrails.” Everyone else waits: developers, enterprises and consumers get it, in the post’s words, “as soon as possible,” starting with paid API customers and Google AI Ultra subscribers.

A gate like that is a claim about capability. Nobody withholds a model from the public and hands it to a vetted set of defenders (the program has 650-plus partners) unless they think it does something the public shouldn’t have yet. So I read what the launch materials let an outsider check: the blog post, the Fairwind program page, and the evaluation PDF that the post links to. I wanted to see how much of the gate’s justification was on the page.

the table

The PDF’s results table has 19 rows, each comparing Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. By my count Argon is first or tied for first in 14 and trails another model in five (FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science 0.1 and OSWorld-2.0). That is a strong table. It is also a table about almost everything except the reason for the gate.

A results table with four columns of scores for Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5, grouped into knowledge work, agentic coding, ML engineering, science and math, long context, computer use, multimodal understanding and cybersecurity. The Argon column is shaded blue where it leads. The last row, CWE-bench v1, reads 68.0%, 68.0%, 58.0% and 67.0%.
Google's own results table for Argon. The cybersecurity group at the bottom is one row. The shading is Google's; the Argon column is the one outlined. Image: Google DeepMind, "Gemini 4 Argon Model evaluation," results table, source PDF, reproduced for commentary. Cropped from page 5.

The one cyber row is CWE-bench v1, which the blog describes as evaluating “the model’s ability to remediate security vulnerabilities.” The PDF says its scores come from the benchmark’s public leaderboard, not from Google’s own runs. The four numbers are below.

CWE-bench v1 scores: Argon ties GPT-6 Astra at 68.0, Opus 5.5 is at 67.0, Fable 5.1 at 58.0 A lollipop chart on an axis from 55 to 70 percent. Gemini 4 Argon 68.0 percent. GPT-6 Astra 68.0 percent. Claude Opus 5.5 67.0 percent. Claude Fable 5.1 58.0 percent. The top three are within one point of each other; Fable 5.1 is ten points below the top. CWE-BENCH V1 · THE ONLY CYBER ROW IN GOOGLE'S ARGON TABLE · SCORE, % 55606570 Gemini 4 Argon GPT-6 Astra Claude Opus 5.5 Claude Fable 5.1 Gemini 4 Argon 68.0% — Google DeepMind evaluation PDF, 30 Sep 2026 GPT-6 Astra 68.0% — Google DeepMind evaluation PDF, 30 Sep 2026 Claude Opus 5.5 67.0% — Google DeepMind evaluation PDF, 30 Sep 2026 Claude Fable 5.1 58.0% — Google DeepMind evaluation PDF, 30 Sep 2026 68.068.067.058.0 Axis starts at 55. The top three sit within one point; the gap to Fable 5.1 is ten. Source: Google's table, scores as printed.
The one public cyber number Google printed is a tie, one point above Opus 5.5. It says Argon is at the front of a pack. It does not, by itself, say why the pack's front needs a waitlist.

Two other cyber evals appear in the methodology section and nowhere in the table. One is an internal dataset that “tests recall over recent, confirmed historical vulnerabilities” in open-source projects. The other is the Wiz Penetration Testing Benchmark, described as internal. The blog says Argon “outperforms 3.8 Flash Cyber” on the Wiz one and “uncovered a wide range of exposures” on Google’s dataset. Neither comes with a score. So of three cyber evals the materials name, one has a number, and it is a tie with a competitor’s model.

That is not evidence that Argon is weaker at cyber than Google says. Labs run internal evals for good reasons, and a full Frontier Safety Framework report may exist that I did not find. In the materials linked from the launch post I found no separate model card and no statement about whether Argon reached a cyber critical capability level. The fair reading is narrower: the public justification for the gate is an assertion, with the numbers that would support it held back by the same company that holds back the model.

Two smaller things. Google ran Argon at “the highest thinking settings,” and the PDF says results for the other models come from “providers’ self reported numbers unless otherwise mentioned,” which means the columns were not produced under one protocol. And the blog’s AI-generated summary says the model has a “1 million token limit,” while the body says Google is “expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K.” Those are different limits; I found no stated input window.

Nineteen rows in Google's Argon results table by category, showing where Argon leads or ties Filled squares mean Argon is first or tied for first on a benchmark row in Google's table; open squares mean another model leads. Knowledge work, four rows, all filled. Agentic coding, four rows: filled, open, filled, open. ML engineering, one row, open. Science and math, three rows: open, filled, filled. Long context, two rows, both filled. Computer use, two rows: filled, open. Multimodal understanding, two rows, both filled. Cybersecurity, one row, filled, a tie with GPT-6 Astra. Fourteen filled, five open, of nineteen. 19 ROWS IN GOOGLE'S TABLE · FILLED = ARGON FIRST OR TIED · MY COUNT Knowledge work Agentic coding ML engineering Science and math Long context Computer use Multimodal Cybersecurity Vals Index — Argon leads AutomationBench — Argon leads Vals Finance Agent v2 — Argon leads Harvey Legal Agent — Argon leads DeepSWE v1.1 — Argon leads Vibe Code Bench — Argon leads LABBench 2 — Argon leads RiemannBench — Argon leads GraphWalks up to 128k — Argon leads GraphWalks 256k to 1M — Argon leads Agent's Last Exam — Argon leads Chartography — Argon leads LVBench — Argon leads CWE-bench v1 — Argon ties GPT-6 Astra at 68.0 FrontierSWE v2 — GPT-6 Astra leads Terminal-bench 4.0 — Claude Opus 5.5 leads PostTrainBench — Claude Opus 5.5 leads Terminal-Bench Science 0.1 — GPT-6 Astra leads OSWorld-2.0 — GPT-6 Astra leads 4 of 4 2 of 4 0 of 1 2 of 3 2 of 2 1 of 2 2 of 2 1 of 1 — a tie with GPT-6 Astra, the only cyber score published Fourteen filled, five open. Rows are as printed in Google's table; the grouping is Google's.
The gate is justified by cyber capability. Of nineteen published rows, eighteen are about something else, and they are where most of Argon's leads are.

who is inside

The Fairwind Program is not new with Argon. Google launched it on 2 September with Gemini 3.8 Flash Cyber and a code-security agent called CodeMender, and said then it had “more than 650 participating partners globally.” Argon was added four weeks later.

The program page is more specific than the launch post about who gets in. Priority goes to “governments and partners most critical to societal resilience”: national cyber authorities, critical-infrastructure operators (healthcare, telecom, energy, finance) and “core technology platforms.” Applicants get background checks. Partners agree to phishing-resistant MFA, may give access only to internal security, incident-response or penetration-testing teams, and are limited to “authorized threat simulation, reverse engineering, and malware analysis for defensive and academic research purposes.” Academic labs “that focus on defensive benchmarking” can apply.

That is a concrete eligibility policy, more than I expected. What it leaves out is who decides. The page says “we will review and respond to eligible partners who meet our criteria as soon as we can.” It gives no timeline, no named review body and no route to contest a rejection. That is ordinary for a private program. It would be unremarkable except that these programs have begun to go wrong in public. In August, TechCrunch reported that five researchers, all outside the US and Europe, said they had lost access to OpenAI’s Trusted Access for Cyber program. One was told by email that access to Daybreak Blue, the latest vetted tier, was revoked “due to a technical issue affecting a limited number of users,” and OpenAI confirmed to TechCrunch that an error was the cause. Affected users were asked to reapply or re-verify. That is a different company and a different program. But it is the one time I found the decision process visible from outside, and it was visible because it broke.

what opens the door

Before broad rollout, the post says Google is “continuing to strengthen critical frontier safeguards across four main areas”: defending against misuse (including cyber and CBRN requests), defending against prompt injection, monitoring for misalignment, and hardening the sandboxes used in training. I went through each for a number.

Misuse: safeguards “underwent robustness testing by internal and external red teams,” with no results given. Prompt injection: Argon is “leading” on the Gray Swan indirect-prompt-injection benchmark, with no figure in the post or the PDF table. Misalignment: a monitor on the model’s reasoning and actions that “stop[s] execution when necessary,” with no rate. Hardening: a roadmap reference. The one external process named is that Google is “actively engaged in the U.S. government’s voluntary process for pre-release model access,” with no stated pass criterion and no agency named.

So the stated preconditions for opening the door are four jobs, none with a threshold, plus a government process with no stated outcome, on a horizon of “as soon as possible.” I have made this complaint twice this month about other documents. The White House accord named two checker roles and no dates, and OpenAI’s training pause in thirty-minutes-to-stop named a resume condition without a pass test. Argon’s is the same shape, though the program page is more specific than either.

There is older background worth a line. Four months ago, when Anthropic widened its Project Glasswing cyber program on 2 June, 9to5Mac reported about 150 new organizations, bringing the total to roughly 200, each of which “will need to meet our security requirements before they gain access.” The article also quoted Anthropic, from its Opus 4.8 introduction a week earlier: “Models of this capability level require stronger cyber safeguards before they can be generally released,” with general availability expected “in the coming weeks.” That is also a condition with no test. It does name the thing (cyber safeguards) and a horizon (weeks), which is a little more checkable than “iterate on guardrails.” I did not check whether the weeks held, and the quotation comes through a news report, not Anthropic’s own page, so I treat it as context and not as evidence.

Staged release itself is not a new idea. Solaiman et al., in a report dated 24 August 2019 on OpenAI’s staged GPT-2 release, described the method as allowing “time between model releases to conduct risk and benefit analyses as model sizes increased.” The gap between stages was supposed to be spent on published analysis. Seven years on, the gap is a waitlist, and the analysis is the part I cannot find: I found no third-party evaluation of Argon’s safety, and the cyber numbers that would show what the gap is for are two internal evals with no scores.

What I will watch for is small. The next Argon document could print a score for either internal cyber eval, name who reviews Fairwind applications, or say what the government process has to conclude before the waitlist ends. Any one of those would turn a posture into a test. Until then, a reader can verify that Argon ties for first on a public leaderboard and has to take the rest of the gate on Google’s word.

References

  1. Kavukcuoglu, K. (2026). Gemini 4 Argon: our next era of frontier intelligence. Google blog, 30 September 2026.
  2. Google DeepMind (2026). Gemini 4 Argon Model evaluation: approach, methodology and results. PDF, published with the launch, 30 September 2026. The results table is an image on page 5; the PDF dates the results “as of October, 2026” in one place and lists benchmarks “as of September, 2026” in another.
  3. Google DeepMind (2026). Fairwind Program. Program page, read 30 September 2026 (undated).
  4. Google (2026). Fairwind Program launch post. Google blog, 2 September 2026.
  5. TechCrunch (2026). Researchers complain that OpenAI revoked their access to limited cyber program. 19 August 2026.
  6. Hall, Z. (2026). Anthropic expands Glasswing as it promises public Claude Mythos-class model releases. 9to5Mac, 2 June 2026. Older than the 90-day news window; used as dated background only. Quotes Anthropic’s statement from its Opus 4.8 introduction, a week earlier; that primary page was not read.
  7. Solaiman, I., Brundage, M., Clark, J., et al. (2019). Release Strategies and the Social Impacts of Language Models. arXiv:1908.09203, 24 August 2019. Historical background.
  8. Gautam Parab (2026). Four Layers of Controls, Zero Deadlines: Reading the White House AI Accord. 30 September 2026.
  9. Gautam Parab (2026). Thirty Minutes to Stop. No Date to Start.. 28 September 2026.