Google released Gemini 4 Argon on 30 September, the first model in the Gemini 4 family and its first new flagship since the Gemini 3 series arrived in November 2025. Almost nobody can use it yet.

The model goes first to a closed group Google calls the Fairwind Program — cyber defenders the company considers trusted — and to Google’s own internal teams. Those two groups get a version of Argon with its cyber guardrails removed. Everyone else waits. Google says developers, enterprises and consumers follow later, starting with paid API customers and Google AI Ultra subscribers, while it strengthens safeguards against cyber and CBRN misuse, indirect prompt injection, model misalignment and insecure agent environments.

The security claim is the headline

Google says Argon can locate a critical software vulnerability, confirm that it is real and write the patch without a person in the loop. It also says the model found a critical flaw in healthcare software used by hospitals worldwide that earlier frontier models had missed. Google has not named the software or published the advisory.

The company is also putting Argon through the US government’s voluntary pre-release model access process, the same arrangement that has been contested elsewhere this month.

Google’s own benchmark table

Google published 18 benchmark comparisons against GPT-6 Astra and Claude Opus 5.5, and Argon leads or ties 13 of them. These are vendor-reported figures, not independent results.

A network patch panel with coloured ethernet cables
The first group to get Argon is a closed set of cyber defenders. Illustrative photo. Brett Sayles · pexels · Pexels License

On DeepSWE v1.1, a software engineering test, Google puts Argon at 77.9% against 74.2% for Opus 5.5. On AutomationBench, which measures business task execution, it reports 51.3% against 42.5%. On LVBench, a long video understanding test, it claims 91.7%. On Harvey’s legal agent evaluation the gap is the widest Google shows: 19.6% for Argon against 5.4% for GPT-6 Astra. On the Gray Swan indirect prompt-injection benchmark Google reports a 0.7% attack success rate, against 8.5% for Astra and 1.0% for Opus 5.5.

The gaps run the other way too, and Google published those. Astra leads FrontierSWE v2 at 65.5% to Argon’s 55.0% and Terminal-Bench Science 0.1 at 68.1% to 57.6%. Opus 5.5 leads Terminal-bench 4.0 at 66.4% to 57.4%.

A million tokens of output

The change most likely to matter in practice is the output limit. Argon can write up to one million output tokens in a single response, up from 64,000 on Google’s previous models. For an agent asked to migrate a codebase or review a long contract, that removes the need to stitch work together across calls.

An empty white corridor lined with doors
Google says Argon found a critical flaw in healthcare software used by hospitals worldwide. Illustrative photo. Enrique Silva · pexels · Pexels License

Google says thousands of its own employees are already using the model, and gives three internal examples: a quantum computing optimisation that beat published baselines by 40%, memory optimisations that freed more than 300 TiB in its data centres, and a rewrite of SIMD code in the libgav1 video decoder that runs 2.7 times faster than the Rust port.

What it costs

Argon launches at an introductory $2 per million input tokens and $10 per million output tokens, with cached input discounted 95% to $0.10. After the introductory period those rise to $4 and $20. At the introductory rate Argon undercuts GPT-6 Astra, which Google lists at $10 and $50, and matches Opus 5.5 once the discount ends.

What to watch

Whether Argon reaches general availability before the end of the year, and whether Google says who is in the Fairwind Program. A model released deliberately without cyber guardrails to an unnamed group of outside organisations is a governance arrangement, not just a product decision, and so far the membership is Google’s to define.