Google has introduced Gemini 4 Argon, a new frontier AI model designed for long-running tasks. These include software engineering, cybersecurity, research and other complex knowledge work.
Announced on September 30, 2026, Gemini 4 Argon is initially being offered to a limited group of cybersecurity defenders through Google DeepMind’s Fairwind Program. Meanwhile, Google says wider access will begin with paid API customers and Google AI Ultra subscribers.
The model stands out for its ability to generate up to 1 million output tokens, a major increase from the previous 64,000-token limit. In particular, that larger output capacity is designed for tasks that require an AI system to work through lengthy processes. It goes beyond simply returning a short response.
Gemini 4 Argon Is Built for Long-Horizon AI Work
The defining idea behind Gemini 4 Argon is sustained work. Google is positioning the model for problems where an AI system needs to examine information, make decisions and test approaches. In addition, the model will continue working across many steps. That includes large software projects, research assignments, business analysis and cybersecurity investigations.
Google reports that Gemini 4 Argon achieved 77.9% on DeepSWE v1.1, a benchmark focused on long-horizon software engineering tasks. The company also reports a 51.3% result on AutomationBench, which evaluates end-to-end task execution across business functions.
These figures are Google’s published benchmark results and represent controlled evaluations. It is important to note that these are not a guarantee of performance in every production environment.
Google Is Already Using Argon for Engineering Work
Google is putting Gemini 4 Argon through internal workloads rather than treating it solely as a research model. The company says thousands of Googlers are already using Argon for coding, research and writing tasks. These projects include work involving complex engineering systems.
One example involves Google’s quantum computing research. Google says Argon helped researchers optimize algorithms and beat a published baseline by 40% in one experiment. Notably, the work was completed in minutes.
The company has also used Argon agents to examine memory efficiency across its data centers. In this case, Google says the agents identified optimization opportunities that could eventually free more than 300 TiB of memory. Estimated savings from these optimizations range from 500 TiB to 1 PiB once the changes are deployed.
Argon Is Being Used to Tackle Large Codebases
Large-scale code migration is another area where Google says Gemini 4 Argon is being tested. Unlike older models generating isolated snippets, the model can work through much larger software repositories. Thus, it helps engineers make changes across complex systems.
Google says Argon agents have been helping migrate C and C++ codebases to Rust. The projects range from smaller core libraries to more than 800,000 lines of code in the Fuchsia Zircon kernel.
Google says these migrations still require automated testing, manual audits, emulation testing and human review before production deployment. This approach reflects an important distinction: the model can perform substantial engineering work, but Google is not removing engineers from the process.
Gemini 4 Argon Has Also Been Tested on Code Optimization
Google has provided another example involving libgav1, its open-source AV1 video decoder. In this project, Argon agents worked on an existing Rust implementation and used repeated profiling and compiler analysis to optimize the code.
According to Google, the agents replaced roughly 32,000 lines of SIMD code and produced a memory-safe decoder that ran 2.7 times faster than the Rust port, while producing identical video output.
The example is notable because the model was not simply asked to write new code. It had to investigate performance, run experiments, interpret the results and repeatedly modify the implementation.
Google Is Expanding Argon Beyond Software Engineering
Coding may be one of the clearest demonstrations of Gemini 4 Argon’s capabilities, but Google is targeting a much broader group of professional workloads. The company says the model can handle tasks across finance, legal services, tax, research and other areas where large amounts of information need to be processed.
Google’s evaluations include the Vals Index, which measures economic impact across professional knowledge-work tasks. The company also reports results from financial and legal research evaluations.
Gemini 4 Argon is multimodal as well, meaning it can work with different types of information rather than relying exclusively on text. In addition, Google reports a 91.7% score on LVBench, a benchmark focused on long-video understanding.
Cybersecurity Is a Major Focus of the Argon Launch
Google is giving Gemini 4 Argon an unusually prominent role in cybersecurity. The company says the model can identify, validate and patch software vulnerabilities. As a result, this creates potential applications for security teams that need to investigate large amounts of code quickly.
The initial cybersecurity rollout is deliberately restricted. Through the Fairwind Program, Google is giving trusted defenders access to the model. They can then use its capabilities in defensive cybersecurity work.
Google says Wiz is testing Argon through its Scan for Good initiative, where AI is being used to identify and remediate security exposures affecting critical infrastructure. Furthermore, Google reports that Argon identified a critical vulnerability in healthcare software that exposed sensitive personal information.
Google Reports Strong Results in Vulnerability Remediation
The cybersecurity testing goes beyond vulnerability discovery. Google is also evaluating whether Gemini 4 Argon can actually help fix the problems it finds.
On CWE-bench v1, Google reports that Gemini 4 Argon tied for the highest score at 68% for vulnerability remediation. The benchmark measures an AI system’s ability to address known software vulnerabilities.
For security teams, that distinction matters. Finding a vulnerability is one task. However, understanding the affected code, developing a workable patch and verifying that the fix does not introduce another problem is considerably more involved.
Google Is Adding More Safeguards Around Agentic AI
Gemini 4 Argon’s expanded capabilities also create additional safety challenges. A model that can work through long tasks and interact with external information has more opportunities to encounter malicious instructions or unexpected conditions.
Google says it has focused on four areas: preventing malicious use, defending against prompt injection, monitoring for potential misalignment and strengthening the environments used for frontier-model testing.
The company says Argon underwent automated red teaming and adversarial training, including testing against indirect prompt injection attacks. Moreover, Google says it uses monitoring systems designed to detect when a model’s actions appear to move beyond the user’s intended objective.
Gemini 4 Argon Will Roll Out Gradually
Google is not opening Gemini 4 Argon to everyone at once. The initial deployment is focused on trusted cybersecurity defenders and testers through the Fairwind Program. This strategy gives Google more opportunity to collect feedback before broader access.
The company says paid API customers and Google AI Ultra subscribers will receive access as the rollout expands.
Google has announced introductory API pricing of $2 per 1 million input tokens and $10 per 1 million output tokens. The company says those rates will later increase to $4 per million input tokens and $20 per million output tokens.
Gemini 4 Argon Points Toward Google’s Agentic AI Strategy
The launch says a lot about where Google sees frontier AI heading. The focus is moving away from models that simply answer questions. Instead, it is moving toward systems capable of carrying out extended pieces of work.
Gemini 4 Argon is being positioned around exactly that idea. It can work through large codebases, conduct research, analyze professional information and assist with cybersecurity tasks that require multiple steps.
The million-token output capacity is part of the same strategy. More output does not automatically mean better AI. However, it gives an agent considerably more room to maintain a long-running task.
For developers and enterprises, the more important question will be what happens when Gemini 4 Argon moves from Google’s controlled evaluations into messy production environments.
That is where the model’s ability to remain accurate, useful and controllable over long tasks will face its biggest test.
Sources
- Google Blog: Gemini 4 Argon
- Google DeepMind: Fairwind Program
- Google DeepMind: Gemini Models

