Google Gemini 4 Argon Matches GPT‑6 Astra in Performance, Cuts Cost by 40%
Google's new flagship model, Gemini 4 Argon, matches GPT-6 Astra on core benchmarks with a 40 percent cost reduction, though it is currently limited to vetted cybersecurity teams.
- Core Development: Google's new flagship model, Gemini 4 Argon, matches GPT-6 Astra on core benchmarks with a 40 percent cost reduction, though it is currently limited to vetted cybersecurity teams.
- Beat Context: Categorized under Business with independent corroboration.
- Reporting Depth: 4 minute analytical read synthesized from verified newsroom sources.
Google has launched Gemini 4 Argon, a new flagship model that ties with OpenAI’s GPT‑6 Astra on several high‑profile benchmarks while promising a 40 % cost reduction for the same level of performance. The model is, however, being released only to vetted cybersecurity teams under the company’s Fairwind Program, leaving developers, enterprise customers and the public on hold.
Performance that rivals the competition
According to Eu.36kr, Gemini 4 Argon outperforms models including GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 in DeepSWE v1.1 test that measures long-range software engineering capabilities, Vals Index that evaluates enterprise knowledge work covering finance, programming, law and taxation, AutomationBench that tests enterprise end-to-end automated execution capabilities, and LVBench that examines long video understanding capabilities; in CWE-bench v1 that assesses vulnerability remediation capabilities, Argon ranks first tied with GPT-6 Astra with a score of 68%. In the LVBench long‑video‑understanding assessment, Argon reached 91.7 %, matching the best existing models. The new model’s 1 million‑token output limit, up from 64,000 tokens in Gemini 3, enables deep, multi‑step reasoning across large codebases and research documents.
Media additions
| Benchmark | Gemini 4 Argon | GPT‑6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1 | 77.9 % | 74.1 % | 74.2 % |
| Vals Index | 68.9 % | 63.1 % | 67.0 % |
| AutomationBench | 51.3 % | — | 42.5 % |
| CWE‑bench | 68 % | 68 % | — |
| LVBench | 91.7 % | , | , |
Cost efficiency that could reshape enterprise budgets
In the Text Arena cost‑performance comparison, Argon entered the frontier with 1525 points and a price of $8 per million tokens. The initial API pricing, as reported by Note, set the cost at $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at $0.1 per million. Artificial Analysis calculates the cost per task to be $1.99, a 40 % reduction over GPT‑6 Astra’s maximum price.
| Metric | Gemini 4 Argon | GPT‑6 Astra |
|---|---|---|
| Cost per task | $1.99 | , |
| Input price (per million tokens) | $2 | , |
| Output price (per million tokens) | $10 | , |
Why the rollout is being held back
Google has restricted Argon’s availability to a vetted group of cybersecurity defenders under the Fairwind Program, as noted by Moneycontrol. The company says the model can autonomously discover, validate and patch critical software vulnerabilities, but the same capabilities could be used to exploit them. The Fairwind Program, described by Thesun.my, is a controlled‑access arrangement for defenders, governments and critical‑infrastructure operators.
Koray Kavukcuoglu, Google DeepMind’s chief AI architect, explained in a blog post that the phased rollout allows the company to gather early feedback, strengthen safety guardrails and avoid misuse. The U.S. Government has been granted early access through a voluntary pre‑release process, according to Straitstimes. The same cautious approach mirrors Anthropic’s handling of its Claude Mythos Preview model.
Safety and guardrails in focus
Google says Argon incorporates four key safety capabilities: abuse prevention, chemical/biological/nuclear weapon restriction, monitoring of internal activation state to detect misalignment, and indirect prompt injection defense. The model rejects requests that could facilitate cyberattacks and is reinforced by automated red‑team testing and adversarial training, as reported by CNBC.
Early testers, including the cybersecurity firm Wiz, have already used Argon to uncover a high‑risk vulnerability in hospital software used worldwide, a flaw that earlier models missed. The company claims the model discovered and remediated the issue autonomously, demonstrating its potential for defensive use.
Internal use cases that showcase Argon’s power
Within Google, Argon has already been employed to optimise memory usage across data centres, freeing hundreds of terabytes without new hardware, and to accelerate quantum‑computing research. The model also assisted in a large‑scale code migration from C/C++ to Rust, rewriting tens of thousands of lines of core library code and over 800,000 lines of the Fuchsia Zircon kernel. In a separate experiment, Argon reduced the execution time of a SIMD‑heavy video decoder by 2.7× while maintaining output consistency.
What comes next?
Google plans to open Argon to paid API customers and Google AI Ultra subscribers after the phased security testing is complete. The company has not yet announced a public release date. In the meantime, developers can monitor the model’s progress through the Fairwind Program, and enterprises can compare Argon’s benchmark scores with those of GPT‑6 Astra and Claude Opus 5.5 to assess potential adoption.
What to watch next
Industry analysts will be comparing Argon’s performance and cost against the latest releases from OpenAI and Anthropic as soon as the model becomes broadly available. The next major milestone will be the first public API pricing announcement, which will determine whether Google can maintain its value proposition in a market that is increasingly focused on cost efficiency.
For the business community, the key question remains whether Argon’s 40 % cost advantage and superior benchmark results will translate into real‑world productivity gains for software engineering, legal research, finance and cyber defence. The outcome will shape the competitive landscape for enterprise AI services in the coming years.
How significant is this development?
Contribute your assessment to the aggregated reader sentiment ledger.
Frequently Asked Questions
Key questions answered in this reportWhat is the key development in: Google Gemini 4 Argon Matches GPT‑6 Astra in Performance, Cuts Cost by 40%?
Google's new flagship model, Gemini 4 Argon, matches GPT-6 Astra on core benchmarks with a 40 percent cost reduction, though it is currently limited to vetted cybersecurity teams.
Why is this Business development significant for the UK?
This report covers critical events in our Business beat. Independent reporting monitors related UK statements, regulatory shifts, and public responses as further verified details emerge.
How was this reporting corroborated and verified?
Newsarchy UK compiles and cross-references reporting from primary reporting from apple.com and cross-checked wire reports. All coverage adheres to published editorial standards.
When was this report published?
This briefing was published on October 1, 2026 and is permanently cataloged in the Newsarchy UK Business archives.