Gemini 4 Argon is Google’s new frontier AI model, and for now almost nobody outside Google can use it. Announced on 30 September 2026, it is being tested first by a small group of “trusted cyber defenders” in Google’s Fairwind Program, while the company takes part in the US government’s voluntary process for pre-release model access. Developers, businesses and consumers will follow “as soon as possible”, starting with paid API customers and Google AI Ultra subscribers.
Two claims dominate the launch. The first is output length: Gemini 4 Argon can produce up to one million tokens in a single response, up from 64,000 on earlier Gemini models. The second is coding: Google says the model sets a new state of the art (SOTA) on DeepSWE v1.1, a benchmark of long software engineering tasks, with 77.9%.
This article explains what Google announced, how the million-token output works and what independent testers measured, how strong the coding claim really is, what Gemini 4 Argon costs, why it is going to cyber defenders first, and what developers should test when it reaches them.
Table of contents
- What Google Announced With Gemini 4 Argon
- The 1M Output Tokens in Gemini 4 Argon, Explained
- The Coding SOTA Claim for Gemini 4 Argon
- Gemini 4 Argon Beyond Coding
- Gemini 4 Argon Pricing and Cost per Task
- Why Gemini 4 Argon Is Going to Cyber Defenders First
- How Gemini 4 Argon Ranks Against GPT-6 Astra and Claude
- What Gemini 4 Argon Means for Developers and Businesses
- Gemini 4 Argon FAQ
- References and Further Reading
What Google Announced With Gemini 4 Argon
The announcement came in a blog post by Koray Kavukcuoglu, chief AI architect at Google, who took over day-to-day leadership of Google DeepMind in August. He described the model as built to “sustain deep reasoning across complex, long-horizon workflows” in software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defence.
A limited test, not a launch
“Safely releasing frontier capabilities at this level requires a phased approach,” Kavukcuoglu wrote. Feedback from early testers will shape the guardrails before wider release, and Google has not given a date. Tulsee Doshi, Google’s Gemini model product lead, told CNBC that starting this way “gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible.”
The headline specifications
According to The Decoder and Artificial Analysis, Gemini 4 Argon keeps a one-million-token input context window and accepts text, images, video and audio, but outputs text only. The pricing is set even though access is not.
| Item | Gemini 4 Argon |
|---|---|
| Announced | 30 September 2026 |
| Access now | Fairwind Program cyber defenders, Google internal teams |
| Next in line | Paid API customers and Google AI Ultra subscribers |
| Input context | 1 million tokens |
| Output limit | 1 million tokens, up from 64,000 |
| Inputs and outputs | Text, image, video and audio in; text out |
| Introductory price | $2 input, $10 output per million tokens |
| Standard price | $4 input, $20 output per million tokens |
| Cached input | 95% off the input price |
What Google has not published
Google has not said when wider access will start, when the introductory price will end, or how large the model is. Bloomberg described Gemini 4 Argon as “a very large model”, and large models cost more to serve, which is one reason analysts are watching how long the discount lasts. The launch post presents its benchmark comparison as a chart, so several figures in this article come from that chart as reported by independent outlets, and others from third-party testers.
Why it took this long
The Decoder notes that Gemini 4 Argon is Google’s first frontier model in more than seven months, after Gemini 3.1 Pro in February. Google promised Gemini 3.5 Pro at its I/O conference in May for June; Bloomberg reported a delay in July, and the model was eventually dropped. In the meantime Google shipped faster, cheaper Flash models, including Gemini 3.8 Flash, which we covered when it reached Google Cloud’s Agent Studio.
The 1M Output Tokens in Gemini 4 Argon, Explained
Most attention has gone on the output limit, and it deserves a closer look, because the headline figure and the tested figure are not the same thing.
From 64,000 to one million
Google calls the one-million-token output limit “industry-leading”. “When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go,” Kavukcuoglu wrote. As a rough guide, a million tokens of English is about 750,000 words.
How Long Decode Continuation works
The million-token figure depends on a new Gemini API feature called Long Decode Continuation. According to The Decoder and Artificial Analysis, it pauses a long response and resumes it through follow-up requests, so the model’s reasoning does not hit a request timeout. Artificial Analysis tested Argon with it and reached the full million tokens.
What independent testers measured
Vals AI lists a maximum direct output of about 262,000 tokens, the amount the model produces in one request without continuation. That is still four times the old 64,000 limit. Developers should therefore treat a million tokens as an orchestrated, multi-request output rather than one uninterrupted reply.
Output ceilings, thousands of tokens (our arithmetic: 262K is 26% of 1M; 64K is 6.4%)
What a long output is for
Very long outputs matter less for chat than for agents. A model migrating a code library, writing a long report from many documents, or working through a hard proof can now keep going in one trajectory instead of being stitched together from short calls. The cost is tokens: long outputs are billed at the output rate, which is five times the input rate.
The Coding SOTA Claim for Gemini 4 Argon
Google’s strongest benchmark claim is in software engineering, and it is where the gap between published scores and day-to-day use is most debated.
DeepSWE v1.1
On DeepSWE v1.1, which measures long, real-world software engineering tasks, Gemini 4 Argon scored 77.9%. Reports of Google’s comparison chart put Claude Opus 5.5 at 74.2% and OpenAI’s GPT-6 Astra at 74.1%. By Latent Space’s count, Argon came first on 13 of the 19 benchmarks in Google’s chart against those two models.
DeepSWE v1.1 scores in Google’s comparison, %
Google’s own engineering examples
Google says thousands of its staff already use the model. Teams of Argon agents are migrating C and C++ code to Rust, from core libraries such as re2 and libgav1 to more than 800,000 lines of the Fuchsia Zircon kernel. For libgav1, its video decoder, agents replaced 32,000 lines of SIMD code with safe Rust that the compiler could vectorise, producing a memory-safe decoder 2.7 times faster than the earlier Rust port. Other agents found data-centre memory savings of more than 300 TiB, with 500 TiB to 1 PiB expected in total.
Independent coding results
Independent tests mostly support a strong but not dominant coder. Vals AI reports that Argon built 30 apps perfectly on its Vibe Code Bench, against 25 for Claude Opus 5 and 24 for GPT-6 Astra, and that its Terminal-Bench 4.0 score rose from 19.0% to 57.6%. Artificial Analysis measured 57% on Terminal Bench 4, behind Claude Sonnet 5.5 at 64%, Claude Opus 5.5 at 60% and GPT-6 Astra at 59%. In Arena’s human-rated Code Arena for web development, Argon placed eighth with 1,679 points.
The doubts inside Google
Bloomberg reported that Gemini 4 Argon “performed well on benchmarks” but “does less well when employees actually put it to work”, particularly on coding, with one insider saying it “isn’t particularly adept at front-end design”. Two people told Bloomberg the model appears affected by “benchmaxxing”, tuning for tests rather than useful work. Google told Bloomberg it would be “inaccurate” to say the model underperforms on coding. The Code Arena result, eighth for web development, fits the front-end criticism.
| Coding test | Gemini 4 Argon | Comparison | Source |
|---|---|---|---|
| DeepSWE v1.1 | 77.9%, first | Opus 5.5 74.2%, Astra 74.1% | |
| Vibe Code Bench | 30 apps | Opus 5: 25, Astra: 24 | Vals AI |
| Terminal Bench 4 | 57% | Sonnet 5.5 64%, Opus 5.5 60%, Astra 59% | Artificial Analysis |
| Code Arena: WebDev | 1,679, eighth | Up from 29th for Gemini 3.8 Flash | Arena |
| CWE-bench v1 (fixing flaws) | 68%, tied first | Tied with GPT-6 Astra and Grok 4.7 | Google, CNBC |
Gemini 4 Argon Beyond Coding
Google pitches the model for knowledge work as much as for code, and here the independent results are stronger.
Knowledge work
Gemini 4 Argon tops the Vals Index at 68.9%, the first Gemini model to do so. The index weights finance, coding, legal and tax work by each sector’s share of US GDP. On Zapier’s AutomationBench, Google reports first place at 51.3%, and on Artificial Analysis’s own version, AutomationBench-AA, Argon also ranks first at 77.5%. One critic noted that its reported 19.6% on Harvey’s legal benchmark trails another listed model’s 25.42%, so the lead is not universal.
Writing and preference
In Arena’s Text Arena, where people compare answers blind, Argon took first place with 1,525 points, 20 ahead of Claude Opus 4.6. The Decoder notes it leads Arena’s categories for coding, hard prompts, instruction following and creative writing.
Agents
In Arena’s preliminary Agent Arena, based on about 3,000 sessions, Argon ranked eighth overall but first for steerability, meaning how well it follows a user’s corrections mid-task. On PostTrainBench, a test of how well a model can run the post-training of another model, it scored 45.3%, up from 21.99% for Gemini 3.1 Pro, according to the Latent Space round-up of launch-day results.
Fewer made-up answers
Artificial Analysis found a 15% hallucination rate on its AA-Omniscience test, the lowest among leading models, against 51% for GPT-6 Astra. The trade-off is accuracy: Argon answered 50% of questions correctly against Astra’s 63%. In other words, it says “I don’t know” more often rather than guessing. For business use, that is often the better failure.
Video and documents
Google reports a state-of-the-art 91.7% on LVBench, a long-video understanding test, and says the model can analyse charts and act on a series of documents.
Gemini 4 Argon Pricing and Cost per Task
Price is where Google is most aggressive, at least for now.
List prices
The introductory rate of $2 per million input tokens and $10 per million output tokens is half the standard rate of $4 and $20. Artificial Analysis says the discount runs for at least a month; Google has not set an end date. Cached input is 95% cheaper than fresh input, up from 90% on Gemini 3.8 Flash.
| Model | Input, per 1M tokens | Output, per 1M tokens | Cached input |
|---|---|---|---|
| Gemini 4 Argon, introductory | $2 | $10 | About $0.10 |
| Gemini 4 Argon, standard | $4 | $20 | About $0.20 |
| GPT-6 Astra | $10 | $50 | $1 |
| GPT-6.1 Sol | $2 | $10 | $0.10 |
| Claude Opus 5.5 | $4 | $20 | $0.20 |
The introductory price matches GPT-6.1 Sol, OpenAI’s cheaper model, while Gemini 4 Argon competes with the far more expensive GPT-6 Astra on capability. At standard rates it costs the same per token as Claude Opus 5.5.
Cost per task
Token prices only tell half the story, because models use different numbers of tokens. Artificial Analysis found Argon averages 62,000 output tokens per task on its Intelligence Index, against 27,000 for GPT-6 Astra. At introductory prices a task costs $1.99, which is 60% of Astra’s $3.26. At standard prices it rises to $3.98, about 1.2 times Astra. GPT-6.1 Sol costs about $0.72 per task.
Cost per Intelligence Index task, US dollars (Artificial Analysis)
What that means for budgets
The saving comes from the discount, not from efficiency. Anyone building on Gemini 4 Argon during the introductory period should model costs at the standard price too, and should measure token use on their own workloads. Vals AI found the opposite pattern on its index, with Argon using about a quarter of Claude Sonnet 5.5’s output tokens, which shows how much the answer depends on the task.
Why Gemini 4 Argon Is Going to Cyber Defenders First
The choice of first users says as much about the model as the benchmarks do.
The Fairwind Program
Fairwind is Google’s channel for government and critical-infrastructure trusted testers. Earlier this month it was the only way to reach Gemini 3.8 Flash Cyber, a restricted security model. Trusted defenders and Google’s own teams will receive the model “without cyber guardrails”, so they can use its full ability to find and fix vulnerabilities.
What it found
Google says the security company Wiz used the model through its Scan for Good programme and found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide, a flaw earlier frontier models had missed. On CWE-bench v1, which tests whether a model can fix security weaknesses, Argon tied for first at 68%. CNBC reported that it tied with GPT-6 Astra and Grok 4.7 on cybersecurity benchmarks.
Safeguards before wider release
Google lists four areas it is strengthening before a broad roll-out: refusing cyber and chemical, biological, radiological and nuclear misuse, including monitoring the model’s internal activations; resisting indirect prompt injection; monitoring the model’s chain of thought and actions for misalignment and stopping it when needed; and sealing the sandboxes used for risky training and evaluation. Google also urged other labs to keep model reasoning readable so that monitoring still works.
The political backdrop
The launch came a day after Google’s chief executive, Sundar Pichai, signed a voluntary White House accord on AI safety with other technology leaders, which we analysed in our piece on AI self-regulation. It also came a day after OpenAI’s developer conference, and the same week OpenAI said it would not release GPT-6.1 Astra over deceptive behaviour. A cautious, staged release is now the norm for the most capable models.
How Gemini 4 Argon Ranks Against GPT-6 Astra and Claude
Taken together, the independent results put Argon level with OpenAI’s best, and just behind Anthropic.
The overall index
On the Artificial Analysis Intelligence Index, Gemini 4 Argon scores 53, the same as GPT-6 Astra and Claude Fable 5.1, and one point above GPT-6.1 Sol. Claude Opus 5.5 leads at 58, with Claude Sonnet 5.5 at 56. That is a 23-point jump on Gemini 3.1 Pro Preview, and Artificial Analysis says it returns Google to the top three labs.
| Model | Intelligence Index | Hallucination rate |
|---|---|---|
| Claude Opus 5.5 | 58 | Not reported here |
| Claude Sonnet 5.5 | 56 | Not reported here |
| Gemini 4 Argon (high) | 53 | 15% |
| GPT-6 Astra (max) | 53 | 51% |
| GPT-6.1 Sol (max) | 52 | 54% |
| Gemini 3.1 Pro Preview | 30 | Not reported here |
Where each model wins
GPT-6 Astra is more accurate on knowledge questions. Claude models lead terminal-based agent work. Argon leads on human preference for text, automation benchmarks, long video and refusing to invent answers. For most teams the right model will depend on the job, which is an argument for testing two or three.
What Gemini 4 Argon Means for Developers and Businesses
Until Google opens access, Gemini 4 Argon is something to plan for rather than build on. The planning is still worth doing now.
When you can use it
Google has given no date. Paid Gemini API customers and Google AI Ultra subscribers come first. Businesses that need it sooner for security work can ask whether they qualify for the Fairwind Program.
What to test when it arrives
Run your own tasks, not the published benchmarks. Check front-end and user interface code in particular, given the Bloomberg report and the Code Arena result. Test long-output jobs with Long Decode Continuation, and confirm how your tools handle a response that arrives across several requests.
How to budget
Price your workloads at both the introductory and standard rates, measure output tokens per task, and use caching wherever prompts repeat, since cached input is 95% cheaper. Compare against Claude Opus 5.5 at the same $4 and $20 list price, and against GPT-6.1 Sol for jobs that do not need a frontier model.
Where it fits for UK organisations
Gemini models reach UK businesses through the Gemini API, Google Cloud and Google Workspace. Before any personal data goes near a new model, check the data processing terms and region settings for the product you use, and record the decision.
Gemini 4 Argon FAQ
What is Gemini 4 Argon?
It is Google’s new frontier AI model, announced on 30 September 2026, aimed at software engineering, knowledge work such as finance and law, and cybersecurity defence.
Can I use Gemini 4 Argon now?
Not unless you are in Google’s Fairwind Program for trusted cyber defenders, or work at Google. Paid API customers and Google AI Ultra subscribers are next, with no date given.
Does Gemini 4 Argon really output one million tokens?
Yes, with the new Long Decode Continuation feature, which resumes a long response across follow-up requests. In a single request, Vals AI measured a maximum of about 262,000 tokens.
Is Gemini 4 Argon the best coding model?
It holds the top published score on DeepSWE v1.1 at 77.9%, but independent tests put it behind Claude and GPT-6 Astra on Terminal Bench 4, and eighth for web development in Code Arena.
How much does Gemini 4 Argon cost?
$2 per million input tokens and $10 per million output tokens at the introductory price, rising to $4 and $20. Cached input is 95% cheaper.
Why is Google limiting access?
Google says a model this capable needs a phased release while it strengthens safeguards against cyber misuse, prompt injection and misaligned behaviour.
References and Further Reading
Gemini 4 Argon: our next era of frontier intelligence (Google)
Google unveils Gemini 4 Argon with SOTA score on DeepSWE (TestingCatalog)
Gemini 4 Argon closes the gap with OpenAI and Anthropic (The Decoder)
Gemini 4 Argon: Google is back as one of the top three labs (Artificial Analysis)
Gemini 4 Argon: GDM’s answer to Astra and Fable, with 1M output (Latent Space)
Google rolls out Gemini 4 Argon, its most advanced AI model (CNBC)
Google announces Gemini 4 for trusted cyber defenders (The Verge)
Gemini 4 and the Bloomberg report on staff doubts (ZeroHedge)
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.