← all posts
// models · gemini

Gemini 4 Argon: 77.9% on DeepSWE, but only for cyber defenders

I wanted to see what Google actually put on its own page before reading anyone's summary of Gemini 4 Argon, and the first thing I noticed is how little of the announcement is about the benchmark everyone quotes. The headline number is 77.9% on DeepSWE v1.1. The operational news is that you cannot use the model yet unless you are a trusted cyber defender.

What Google states on its own page

Google announced Argon on September 30, 2026. It is rolling out first to "trusted cyber defenders" through the Fairwind Program, while a U.S. government voluntary pre-release process is underway. Broader availability for developers, enterprises and consumers is promised "as soon as possible", with no date attached. That is the entire access story, and I would not read more into it than that.

The numbers Google lists: 77.9% on DeepSWE v1.1, 51.3% on Zapier's AutomationBench (which Google presents as first place), 68% on CWE-bench v1 (a tie for first), 91.7% on LVBench for long video understanding, and leading resilience on the Gray Swan indirect prompt injection benchmark. It also claims leading results on the Vals Index across finance, coding, legal and tax. The output limit jumps to 1M tokens, up from 64K on the previous model.

Press coverage of the launch puts Claude Opus 5.5 at 74.2% and GPT-6 Astra at 74.1% on DeepSWE v1.1, which would make the gap about 3.7 points. I could not confirm those two figures on Google's page, so treat them as reported, not verified.

The price line with an expiry date

A launch price with no end date is a loan, and the interest rate is 2x.

Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input at a 95% discount. Google also states that $4 and $20 will apply after the introductory period. It does not say when that period ends.

Here is a small illustrative calculation, with assumptions I made up for the purpose: an agent run that reads 2M input tokens (none cached) and writes 100K output tokens. At intro pricing that is 2 x $2 + 0.1 x $10 = $5.00. At the stated post-intro rates it is 2 x $4 + 0.1 x $20 = $10.00. The run costs exactly double, and any budget you build on the $2/$10 numbers is wrong by a factor of two the day the period ends.

The cache discount matters more than the headline rate for coding agents, because most of their input is re-read context (see where agent loops leak tokens). If 90% of those 2M tokens were cache hits, the arithmetic changes a lot, but I do not know the cache storage terms or how Google bills them, so I will not invent a number.

A 1M output window is a different kind of change

Most of the attention goes to the context size, but the jump from 64K to 1M output tokens changes what a single call can be. Google's examples include C/C++ to Rust migrations of codebases up to 800K+ lines. I read that as a claim about a model that can emit a whole translated module set in one pass, not a chat reply.

I am skeptical, and the reason is practical. A single 1M token generation is a long time to wait, a long time to pay for, and an enormous diff to review. If it goes wrong at token 400K you have bought a lot of wrong. My guess is that real users will still chunk the work, and the large limit will mostly buy them freedom from truncation errors. That is a guess, not something Google claims.

What I would do before the model is available

Nothing in the announcement lets me test Argon, and I have not run it. So the useful work now is preparation. If you run a coding agent on DeepSWE-like tasks, pin a baseline on your current model with your own repositories, because a vendor benchmark of long-horizon tasks tells you little about your build system and your test flakiness. If you already model your agent spend, add a second column at double the intro price. And if you care about prompt injection (you should, for any agent that reads web pages or tickets), note that Gray Swan is the only prompt injection benchmark Google cites, and "leading" without a number is not something I can plan around.

There is also the cyber angle. Gating the first release to defenders is a deliberate choice, and it fits the CWE-bench result: a model good at finding and fixing vulnerabilities is good at the other half of that job too. How long the gate stays up is the open question, and it will probably be decided in the pre-release process Google mentions rather than by the benchmark tables.

I will revisit this when there is a public endpoint and a price I can actually be billed.

#gemini#coding-agents#pricing