Quick Summary
- Google announced Gemini 4 Argon on September 30, 2026, as a frontier model for coding, professional knowledge work, and cyber defense.
- Access starts with trusted cyber defenders through the Fairwind Program.
- The output limit rises to 1 million tokens, up from 64,000.
- Google reports 77.9% on DeepSWE v1.1 and a first place rank on AutomationBench at 51.3%.
- Google teams use Argon for quantum research, data center memory savings, and Rust migrations.
- Argon ties for first on CWE-bench v1 with 68%, a test of how well models fix security flaws.
- Launch pricing is $2 per million input tokens and $10 per million output tokens.
What Google Announced With Gemini 4 Argon
Gemini 4 Argon is Google’s newest frontier AI model, announced on September 30, 2026. Google built it to handle long and complex jobs in software engineering, professional knowledge work, and cybersecurity defense. This guide covers what the model does, how Google says it performs, and when people can expect to use it.
Access is limited for now. Argon is rolling out first to a group of trusted cyber defenders through Google’s Fairwind Program. Google calls its approach phased and says it is taking part in the U.S. government’s voluntary process for pre-release model access.
Launch pricing is set at $2 per million input tokens and $10 per million output tokens. Cached input tokens cost 95% less than the standard input price. Google says the rates move to $4 and $20 per million tokens once the introductory period ends.
How Google Is Putting Argon to Work Inside Its Own Teams
Google says thousands of its employees already use Argon in daily work. They point to better results in specialized coding, deeper research, and writing quality. The company shares a few examples of what the model has done so far.
Google’s quantum computing researchers used Argon to make a key step in a quantum program cheaper to run. Google counts that cost in qubits, the basic units of a quantum computer, and gates, the operations performed on them. In one test the model beat a published reference result by 40% in a matter of minutes.
A team of Argon agents also studied performance data from across Google’s data centers and applied ways to use memory more efficiently. Google says this will free up more than 300 TiB of memory once rolled out. Total savings are estimated at 500 TiB to 1 PiB. For scale, 1 TiB is roughly 1,000 gigabytes.
Argon agents are also helping rewrite older C and C++ code in Rust, a newer programming language built to prevent many memory-related bugs. The projects range from libraries with tens of thousands of lines of code to the Fuchsia Zircon kernel, the core of an operating system, at more than 800,000 lines. Google says these rewrites are going through automated checks, manual audits, and testing before they reach production.
The libgav1 project, Google’s open source software for decoding video, shows how this works. Argon agents took an existing Rust version and replaced 32,000 lines of speed-focused code with safer code. They tested many changes and studied how the compiler turns code into instructions, which let them write safe Rust that the compiler can speed up on its own. The result is a video decoder that is safe with memory and runs 2.7 times faster than the earlier Rust version, with identical video output.
From 64,000 to 1 Million Tokens of Output
Tokens are small chunks of text that a model reads and writes. Google raised Argon’s output limit to 1 million tokens, up from 64,000. The company calls this an industry-leading figure.
The larger limit gives the model room to think at length. It can produce hundreds of thousands of tokens in a single run. Google says this adds depth when the model works on tough problems in one pass.
Benchmark Results for Coding and Business Work
Google reports a new state-of-the-art result on DeepSWE v1.1, where Argon scores 77.9%. The test measures how well a model handles real software engineering tasks that run for a long time. Google engineers use the model for everyday debugging, large codebase migrations, and algorithm design.
Argon also leads the Vals Index, which measures economic impact across finance, coding, legal, and tax work. Each sector is weighted by its share of U.S. GDP.
Google reports similar leading results on Vals Finance Agent v2 for multi-step financial research and on Harvey’s Legal Agent Benchmark for legal research and drafting. On Zapier’s AutomationBench, a test of end-to-end business tasks, Argon ranks first with 51.3%.
Google also highlights visual understanding. The model can analyze professional charts, pick out details in long videos, and act on information spread across several documents. On LVBench, a long video understanding test, it scored 91.7%, which Google lists as state of the art. Independent testing will matter once the model is widely available because these scores come from Google’s own reporting.
Finding and Patching Security Flaws on Its Own
Google trained Argon to be highly capable at cybersecurity defense. The model can find, validate, and patch critical software vulnerabilities without step-by-step human direction. Trusted defenders and Google’s internal teams will receive a version without cyber guardrails so they can use its full defensive power.
Security firm Wiz already uses Argon through its Scan for Good program, which protects critical public infrastructure for free. In an early demonstration, the model found a critical flaw that exposed sensitive personal data in healthcare software used by hospitals worldwide. Google says earlier frontier models had missed this risk.
On CWE-bench v1, which tests how well a model fixes security flaws, Argon ties for first place with 68%. Security teams often sort flaws using the Common Weakness Enumeration, a public catalog of software and hardware weaknesses.
Google also reports gains in vulnerability discovery over its earlier Gemini 3.8 Flash Cyber model. On an internal benchmark, Argon found exposures in codebases that span 20 programming languages. On Wiz’s black-box penetration test, which gives no access to source code, it did better at mapping the attack surface, spotting flaws, and producing proof-of-concept evidence.
The Safety Work Behind a Slow Rollout
Google says it is strengthening safeguards in four areas before a broad launch. The first is misuse. The model is designed to refuse harmful requests tied to cyber attacks or to chemical, biological, radiological, and nuclear threats while still supporting legitimate dual-use science.
Google is improving how it watches the model’s internal activations to spot misuse. Internal and external red teams tested these safeguards with manual and automated attacks.
The second area is prompt injection. In these attacks, hidden instructions in content the model reads can change its behavior, a risk that the OWASP Gen AI Security Project ranks first among large language model threats. Google calls Argon its most resilient model yet against indirect prompt injection. It says the model leads on Gray Swan’s Indirect Prompt Injection benchmark.
The third area is misalignment. Google monitors the model’s chain of thought and actions. It stops execution when the model goes beyond what the user intended. A similar system watched training runs and alerted an incident response team.
Google took care not to feed those findings back into training. This avoids teaching the model’s reasoning to evade the monitoring. The company also urges the wider industry to keep reasoning transparent during this period.
The fourth area is system hardening. Google is isolating and sealing its sandboxed test environments before high-risk training or evaluations begin. It plans to share these agent security practices with partners.
Who Gets Access First, and What Comes Next
Google says Argon will reach developers, enterprises, and consumers as soon as possible. The release starts with paid API customers and Google AI Ultra subscribers. The company thanks the first group of cyber defenders and trusted testers, whose feedback will help tighten the systems before launch.
Readers who plan to build with Argon should watch for independent evaluations after launch. Pricing will also change once the introductory period ends.
Final Thoughts
Gemini 4 Argon is a model built for long, difficult work in coding, business tasks, and cyber defense. Google reports strong benchmark scores, a much larger output limit, and real gains inside its own engineering teams. Access remains limited while the company tests its safeguards with trusted partners.
The wider release will show how these results hold up outside Google. Developers and teams can prepare by learning the pricing and watching for early independent reviews.
Discover how AI is reshaping technology, business, and healthcare—without the hype.
Visit InfluenceOfAI.com for easy-to-understand insights, expert analysis, and real-world applications of artificial intelligence. From the latest tools to emerging trends, we help you navigate the AI landscape with clarity and confidence