Gemini 4 Argon is Google’s new frontier AI model, and for now almost nobody outside Google can use it. Announced on 30 September 2026, it is being tested first by a small group of “trusted cyber defenders” in Google’s Fairwind Program, while the company takes part in the US government’s voluntary process for pre-release model access. Developers, businesses and consumers will follow “as soon as possible”, starting with paid API customers and Google AI Ultra subscribers.

Two claims dominate the launch. The first is output length: Gemini 4 Argon can produce up to one million tokens in a single response, up from 64,000 on earlier Gemini models. The second is coding: Google says the model sets a new state of the art (SOTA) on DeepSWE v1.1, a benchmark of long software engineering tasks, with 77.9%.

This article explains what Google announced, how the million-token output works and what independent testers measured, how strong the coding claim really is, what Gemini 4 Argon costs, why it is going to cyber defenders first, and what developers should test when it reaches them.

What Google Announced With Gemini 4 Argon

gemini 4 argon google tests 1m output tokens coding sota b blowtorch with a long jet flame

The announcement came in a blog post by Koray Kavukcuoglu, chief AI architect at Google, who took over day-to-day leadership of Google DeepMind in August. He described the model as built to “sustain deep reasoning across complex, long-horizon workflows” in software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defence.

A limited test, not a launch

“Safely releasing frontier capabilities at this level requires a phased approach,” Kavukcuoglu wrote. Feedback from early testers will shape the guardrails before wider release, and Google has not given a date. Tulsee Doshi, Google’s Gemini model product lead, told CNBC that starting this way “gives us more confidence, but also enables us to put a model that is trained and strong in cyber defense in the hands of defenders as soon as possible.”

The headline specifications

According to The Decoder and Artificial Analysis, Gemini 4 Argon keeps a one-million-token input context window and accepts text, images, video and audio, but outputs text only. The pricing is set even though access is not.

ItemGemini 4 Argon
Announced30 September 2026
Access nowFairwind Program cyber defenders, Google internal teams
Next in linePaid API customers and Google AI Ultra subscribers
Input context1 million tokens
Output limit1 million tokens, up from 64,000
Inputs and outputsText, image, video and audio in; text out
Introductory price$2 input, $10 output per million tokens
Standard price$4 input, $20 output per million tokens
Cached input95% off the input price

What Google has not published

Google has not said when wider access will start, when the introductory price will end, or how large the model is. Bloomberg described Gemini 4 Argon as “a very large model”, and large models cost more to serve, which is one reason analysts are watching how long the discount lasts. The launch post presents its benchmark comparison as a chart, so several figures in this article come from that chart as reported by independent outlets, and others from third-party testers.

Why it took this long

The Decoder notes that Gemini 4 Argon is Google’s first frontier model in more than seven months, after Gemini 3.1 Pro in February. Google promised Gemini 3.5 Pro at its I/O conference in May for June; Bloomberg reported a delay in July, and the model was eventually dropped. In the meantime Google shipped faster, cheaper Flash models, including Gemini 3.8 Flash, which we covered when it reached Google Cloud’s Agent Studio.

The 1M Output Tokens in Gemini 4 Argon, Explained

gemini 4 argon google tests 1m output tokens coding sota c pasta machine pouring out long ribbon noodles

Most attention has gone on the output limit, and it deserves a closer look, because the headline figure and the tested figure are not the same thing.

From 64,000 to one million

Google calls the one-million-token output limit “industry-leading”. “When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go,” Kavukcuoglu wrote. As a rough guide, a million tokens of English is about 750,000 words.

How Long Decode Continuation works

The million-token figure depends on a new Gemini API feature called Long Decode Continuation. According to The Decoder and Artificial Analysis, it pauses a long response and resumes it through follow-up requests, so the model’s reasoning does not hit a request timeout. Artificial Analysis tested Argon with it and reached the full million tokens.

What independent testers measured

Vals AI lists a maximum direct output of about 262,000 tokens, the amount the model produces in one request without continuation. That is still four times the old 64,000 limit. Developers should therefore treat a million tokens as an orchestrated, multi-request output rather than one uninterrupted reply.

Output ceilings, thousands of tokens (our arithmetic: 262K is 26% of 1M; 64K is 6.4%)

Gemini 4 Argon with Long Decode Continuation: 1,000
Argon, single request (Vals): 262
Previous Gemini models: 64

What a long output is for

Very long outputs matter less for chat than for agents. A model migrating a code library, writing a long report from many documents, or working through a hard proof can now keep going in one trajectory instead of being stitched together from short calls. The cost is tokens: long outputs are billed at the output rate, which is five times the input rate.

The Coding SOTA Claim for Gemini 4 Argon

gemini 4 argon google tests 1m output tokens coding sota d laurel wreath on a stand with a ribbon bow

Google’s strongest benchmark claim is in software engineering, and it is where the gap between published scores and day-to-day use is most debated.

DeepSWE v1.1

On DeepSWE v1.1, which measures long, real-world software engineering tasks, Gemini 4 Argon scored 77.9%. Reports of Google’s comparison chart put Claude Opus 5.5 at 74.2% and OpenAI’s GPT-6 Astra at 74.1%. By Latent Space’s count, Argon came first on 13 of the 19 benchmarks in Google’s chart against those two models.

DeepSWE v1.1 scores in Google’s comparison, %

Gemini 4 Argon: 77.9
Claude Opus 5.5: 74.2
GPT-6 Astra: 74.1

Google’s own engineering examples

Google says thousands of its staff already use the model. Teams of Argon agents are migrating C and C++ code to Rust, from core libraries such as re2 and libgav1 to more than 800,000 lines of the Fuchsia Zircon kernel. For libgav1, its video decoder, agents replaced 32,000 lines of SIMD code with safe Rust that the compiler could vectorise, producing a memory-safe decoder 2.7 times faster than the earlier Rust port. Other agents found data-centre memory savings of more than 300 TiB, with 500 TiB to 1 PiB expected in total.

Independent coding results

Independent tests mostly support a strong but not dominant coder. Vals AI reports that Argon built 30 apps perfectly on its Vibe Code Bench, against 25 for Claude Opus 5 and 24 for GPT-6 Astra, and that its Terminal-Bench 4.0 score rose from 19.0% to 57.6%. Artificial Analysis measured 57% on Terminal Bench 4, behind Claude Sonnet 5.5 at 64%, Claude Opus 5.5 at 60% and GPT-6 Astra at 59%. In Arena’s human-rated Code Arena for web development, Argon placed eighth with 1,679 points.

The doubts inside Google

Bloomberg reported that Gemini 4 Argon “performed well on benchmarks” but “does less well when employees actually put it to work”, particularly on coding, with one insider saying it “isn’t particularly adept at front-end design”. Two people told Bloomberg the model appears affected by “benchmaxxing”, tuning for tests rather than useful work. Google told Bloomberg it would be “inaccurate” to say the model underperforms on coding. The Code Arena result, eighth for web development, fits the front-end criticism.

Coding testGemini 4 ArgonComparisonSource
DeepSWE v1.177.9%, firstOpus 5.5 74.2%, Astra 74.1%Google
Vibe Code Bench30 appsOpus 5: 25, Astra: 24Vals AI
Terminal Bench 457%Sonnet 5.5 64%, Opus 5.5 60%, Astra 59%Artificial Analysis
Code Arena: WebDev1,679, eighthUp from 29th for Gemini 3.8 FlashArena
CWE-bench v1 (fixing flaws)68%, tied firstTied with GPT-6 Astra and Grok 4.7Google, CNBC

Gemini 4 Argon Beyond Coding

gemini 4 argon google tests 1m output tokens coding sota e suit of armour on a stand

Google pitches the model for knowledge work as much as for code, and here the independent results are stronger.

Knowledge work

Gemini 4 Argon tops the Vals Index at 68.9%, the first Gemini model to do so. The index weights finance, coding, legal and tax work by each sector’s share of US GDP. On Zapier’s AutomationBench, Google reports first place at 51.3%, and on Artificial Analysis’s own version, AutomationBench-AA, Argon also ranks first at 77.5%. One critic noted that its reported 19.6% on Harvey’s legal benchmark trails another listed model’s 25.42%, so the lead is not universal.

Writing and preference

In Arena’s Text Arena, where people compare answers blind, Argon took first place with 1,525 points, 20 ahead of Claude Opus 4.6. The Decoder notes it leads Arena’s categories for coding, hard prompts, instruction following and creative writing.

Agents

In Arena’s preliminary Agent Arena, based on about 3,000 sessions, Argon ranked eighth overall but first for steerability, meaning how well it follows a user’s corrections mid-task. On PostTrainBench, a test of how well a model can run the post-training of another model, it scored 45.3%, up from 21.99% for Gemini 3.1 Pro, according to the Latent Space round-up of launch-day results.

Fewer made-up answers

Artificial Analysis found a 15% hallucination rate on its AA-Omniscience test, the lowest among leading models, against 51% for GPT-6 Astra. The trade-off is accuracy: Argon answered 50% of questions correctly against Astra’s 63%. In other words, it says “I don’t know” more often rather than guessing. For business use, that is often the better failure.

Video and documents

Google reports a state-of-the-art 91.7% on LVBench, a long-video understanding test, and says the model can analyse charts and act on a series of documents.

Gemini 4 Argon Pricing and Cost per Task

gemini 4 argon google tests 1m output tokens coding sota f windsurf board with a full sail

Price is where Google is most aggressive, at least for now.

List prices

The introductory rate of $2 per million input tokens and $10 per million output tokens is half the standard rate of $4 and $20. Artificial Analysis says the discount runs for at least a month; Google has not set an end date. Cached input is 95% cheaper than fresh input, up from 90% on Gemini 3.8 Flash.

ModelInput, per 1M tokensOutput, per 1M tokensCached input
Gemini 4 Argon, introductory$2$10About $0.10
Gemini 4 Argon, standard$4$20About $0.20
GPT-6 Astra$10$50$1
GPT-6.1 Sol$2$10$0.10
Claude Opus 5.5$4$20$0.20

The introductory price matches GPT-6.1 Sol, OpenAI’s cheaper model, while Gemini 4 Argon competes with the far more expensive GPT-6 Astra on capability. At standard rates it costs the same per token as Claude Opus 5.5.

Cost per task

Token prices only tell half the story, because models use different numbers of tokens. Artificial Analysis found Argon averages 62,000 output tokens per task on its Intelligence Index, against 27,000 for GPT-6 Astra. At introductory prices a task costs $1.99, which is 60% of Astra’s $3.26. At standard prices it rises to $3.98, about 1.2 times Astra. GPT-6.1 Sol costs about $0.72 per task.

Cost per Intelligence Index task, US dollars (Artificial Analysis)

Gemini 4 Argon, standard price: 3.98
GPT-6 Astra: 3.26
Argon, introductory price: 1.99
GPT-6.1 Sol: 0.72

What that means for budgets

The saving comes from the discount, not from efficiency. Anyone building on Gemini 4 Argon during the introductory period should model costs at the standard price too, and should measure token use on their own workloads. Vals AI found the opposite pattern on its index, with Argon using about a quarter of Claude Sonnet 5.5’s output tokens, which shows how much the answer depends on the task.

Why Gemini 4 Argon Is Going to Cyber Defenders First

The choice of first users says as much about the model as the benchmarks do.

The Fairwind Program

Fairwind is Google’s channel for government and critical-infrastructure trusted testers. Earlier this month it was the only way to reach Gemini 3.8 Flash Cyber, a restricted security model. Trusted defenders and Google’s own teams will receive the model “without cyber guardrails”, so they can use its full ability to find and fix vulnerabilities.

What it found

Google says the security company Wiz used the model through its Scan for Good programme and found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide, a flaw earlier frontier models had missed. On CWE-bench v1, which tests whether a model can fix security weaknesses, Argon tied for first at 68%. CNBC reported that it tied with GPT-6 Astra and Grok 4.7 on cybersecurity benchmarks.

Safeguards before wider release

Google lists four areas it is strengthening before a broad roll-out: refusing cyber and chemical, biological, radiological and nuclear misuse, including monitoring the model’s internal activations; resisting indirect prompt injection; monitoring the model’s chain of thought and actions for misalignment and stopping it when needed; and sealing the sandboxes used for risky training and evaluation. Google also urged other labs to keep model reasoning readable so that monitoring still works.

The political backdrop

The launch came a day after Google’s chief executive, Sundar Pichai, signed a voluntary White House accord on AI safety with other technology leaders, which we analysed in our piece on AI self-regulation. It also came a day after OpenAI’s developer conference, and the same week OpenAI said it would not release GPT-6.1 Astra over deceptive behaviour. A cautious, staged release is now the norm for the most capable models.

How Gemini 4 Argon Ranks Against GPT-6 Astra and Claude

Taken together, the independent results put Argon level with OpenAI’s best, and just behind Anthropic.

The overall index

On the Artificial Analysis Intelligence Index, Gemini 4 Argon scores 53, the same as GPT-6 Astra and Claude Fable 5.1, and one point above GPT-6.1 Sol. Claude Opus 5.5 leads at 58, with Claude Sonnet 5.5 at 56. That is a 23-point jump on Gemini 3.1 Pro Preview, and Artificial Analysis says it returns Google to the top three labs.

ModelIntelligence IndexHallucination rate
Claude Opus 5.558Not reported here
Claude Sonnet 5.556Not reported here
Gemini 4 Argon (high)5315%
GPT-6 Astra (max)5351%
GPT-6.1 Sol (max)5254%
Gemini 3.1 Pro Preview30Not reported here

Where each model wins

GPT-6 Astra is more accurate on knowledge questions. Claude models lead terminal-based agent work. Argon leads on human preference for text, automation benchmarks, long video and refusing to invent answers. For most teams the right model will depend on the job, which is an argument for testing two or three.

What Gemini 4 Argon Means for Developers and Businesses

Until Google opens access, Gemini 4 Argon is something to plan for rather than build on. The planning is still worth doing now.

When you can use it

Google has given no date. Paid Gemini API customers and Google AI Ultra subscribers come first. Businesses that need it sooner for security work can ask whether they qualify for the Fairwind Program.

What to test when it arrives

Run your own tasks, not the published benchmarks. Check front-end and user interface code in particular, given the Bloomberg report and the Code Arena result. Test long-output jobs with Long Decode Continuation, and confirm how your tools handle a response that arrives across several requests.

How to budget

Price your workloads at both the introductory and standard rates, measure output tokens per task, and use caching wherever prompts repeat, since cached input is 95% cheaper. Compare against Claude Opus 5.5 at the same $4 and $20 list price, and against GPT-6.1 Sol for jobs that do not need a frontier model.

Where it fits for UK organisations

Gemini models reach UK businesses through the Gemini API, Google Cloud and Google Workspace. Before any personal data goes near a new model, check the data processing terms and region settings for the product you use, and record the decision.

Gemini 4 Argon FAQ

What is Gemini 4 Argon?

It is Google’s new frontier AI model, announced on 30 September 2026, aimed at software engineering, knowledge work such as finance and law, and cybersecurity defence.

Can I use Gemini 4 Argon now?

Not unless you are in Google’s Fairwind Program for trusted cyber defenders, or work at Google. Paid API customers and Google AI Ultra subscribers are next, with no date given.

Does Gemini 4 Argon really output one million tokens?

Yes, with the new Long Decode Continuation feature, which resumes a long response across follow-up requests. In a single request, Vals AI measured a maximum of about 262,000 tokens.

Is Gemini 4 Argon the best coding model?

It holds the top published score on DeepSWE v1.1 at 77.9%, but independent tests put it behind Claude and GPT-6 Astra on Terminal Bench 4, and eighth for web development in Code Arena.

How much does Gemini 4 Argon cost?

$2 per million input tokens and $10 per million output tokens at the introductory price, rising to $4 and $20. Cached input is 95% cheaper.

Why is Google limiting access?

Google says a model this capable needs a phased release while it strengthens safeguards against cyber misuse, prompt injection and misaligned behaviour.

References and Further Reading