Hero Background

Power Your Software Testing with AI Agents and Cloud

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.

AIAgent Testing

Gemini 4 Argon Explained: Benchmarks, Pricing, and Access

Gemini 4 Argon is Google's new frontier model. See its benchmarks against GPT-6 Astra and Claude Opus 5.5, pricing, 1M-token output limit, and who can use it.

Published on:

Gemini 4 Argon can generate up to 1 million output tokens in a single response, up from the 64K limit of Google's previous models, according to Google's announcement. Google calls it its next frontier model, built to sustain deep reasoning across complex, long-horizon workflows.

Most teams cannot use it yet. Argon is rolling out first to trusted cyber defenders, with paid API customers next and developers, enterprises, and consumers after that.

For teams building on Gemini, Google's benchmark table also shows where Argon trails rival models, and the larger output limit changes how long one agent run can take and what it can cost. TestMu AI checks for coding agents and conversational agents fit in before you switch a production agent over.

TL;DR

Gemini 4 Argon is Google's frontier AI model, announced on 30 September 2026 and built for long-horizon software engineering, legal and finance work, and cybersecurity defense. It can generate up to 1M output tokens per response and is rolling out first to trusted cyber defenders in Google's Fairwind Program.

  • Availability: Can developers use Gemini 4 Argon today? No. It reaches Fairwind Program cyber defenders first, then paid API customers and Google AI Ultra subscribers, before developers, enterprises, and consumers.
  • Pricing: Gemini 4 Argon launches at an introductory $2 per million input tokens and $10 per million output tokens, rising to $4 and $20, with cached input 95% off the input price.
  • Output limit: Is Gemini 4 Argon's output limit larger than earlier Gemini models? Yes. The limit rises to 1M tokens per response, up from 64K in previous Gemini models.
  • Benchmarks: Does Gemini 4 Argon beat Claude Opus 5.5 on every benchmark? No. In Google's table it leads on most rows but trails Claude Opus 5.5 on Terminal-bench 4.0 and PostTrainBench.
  • Testing agents: TestMu AI Kane CLI checks what an Argon-powered coding agent ships in a real browser, and Agent Testing evaluates chat, voice, and phone agents before you move them to a new model.

What Is Gemini 4 Argon?

Gemini 4 Argon is Google's new frontier model, announced on 30 September 2026 by Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect at Google. Google aims it at real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense.

Google says thousands of Googlers already use Argon, and its announcement lists these internal results:

  • Quantum algorithms - optimizing the spacetime resources of subroutines that bottleneck quantum applications, where in one example it beat the published baseline by 40% in minutes.
  • Memory efficiency - a team of Argon agents analyzed fleet-wide profiling telemetry and freed over 300 TiB of memory across Google's data centers, with an estimated 500 TiB to 1 PiB in total savings.
  • C/C++ to Rust migrations - Argon agents are migrating codebases from tens of thousands of lines in libraries such as re2 and libgav1 up to 800K+ lines for the Fuchsia Zircon kernel, with automated and manual auditing before production.
  • Faster safe code - for libgav1, Argon replaced 32K lines of SIMD code in an existing Rust port, producing a memory-safe video decoder that runs 2.7x faster than that port with identical output.
Youtube thumbnail

Gemini 4 Argon Benchmarks: Where It Leads and Where It Trails

Google's announcement compares Argon with OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 and Claude Opus 5.5. These are Google's own reported scores, with the highest score in each row in bold:

BenchmarkGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5
Vals Index (knowledge work)68.9%63.1%65.8%67.0%
AutomationBench51.3%41.4%31.4%42.5%
Vals Finance Agent v265.4%53.5%58.9%58.6%
Harvey's Legal Agent Benchmark19.6%5.4%6.7%3.8%
DeepSWE v1.1 (agentic coding)77.9%74.1%67.4%74.2%
FrontierSWE v255.0%65.5%56.3%62.3%
Vibe Code Bench91.9%89.6%90.3%90.3%
Terminal-bench 4.057.4%58.2%57.9%66.4%
PostTrainBench (ML engineering)45.3%44.3%40.2%49.3%
Terminal-Bench Science 0.157.6%68.1%52.6%63.3%
LABBench 288.8%85.4%68.6%73.1%
RiemannBench76.0%72.0%65.6%69.6%
GraphWalks, up to 128K tokens99.7%98.7%91.4%90.6%
GraphWalks, 256K to 1M tokens84.2%71.8%65.0%66.8%
Agent's Last Exam (computer use, pass rate)39.5%34.2%Not reported38.2%
OSWorld-2.0 (offline subset)69.2%72.6%Not reportedNot reported
Chartography71.6%71.0%46.2%66.3%
LVBench (long video)91.7%87.5%79.7%83.7%
CWE-bench v1 (cybersecurity)68.0%68.0%58.0%67.0%

Argon's widest leads come in knowledge work and very long context:

  • Legal work - 19.6% on Harvey's Legal Agent Benchmark, against 6.7% for the next-best model.
  • Business automation - 51.3% on AutomationBench, Zapier's end-to-end business workflow benchmark, 8.8 points ahead of Claude Opus 5.5.
  • Very long context - 84.2% on GraphWalks between 256K and 1M tokens, against 71.8% for GPT-6 Astra.
  • Finance research - 65.4% on Vals Finance Agent v2, against 58.9% for Claude Fable 5.1.

It trails on several of the coding and computer-use rows that matter most to agent builders:

  • Terminal tasks - 57.4% on Terminal-bench 4.0, behind Claude Opus 5.5 at 66.4%.
  • Frontier software engineering - 55.0% on FrontierSWE v2, behind GPT-6 Astra at 65.5% and Claude Opus 5.5 at 62.3%.
  • Computer use - 69.2% on the OSWorld-2.0 offline subset, behind GPT-6 Astra at 72.6%.
  • Science and ML engineering - behind GPT-6 Astra on Terminal-Bench Science 0.1 and behind Claude Opus 5.5 on PostTrainBench.

Argon tops DeepSWE v1.1, Google's headline long-horizon coding result, yet a terminal-heavy workload may still run better on another model, so compare on your own tasks. For how Anthropic's model behaves in agent work, see Claude Opus 5.5.

What Does the 1M-Token Output Limit Change?

Argon's output limit rises to 1M tokens, up from 64K in Google's previous models. Google's announcement says that with that headroom the model can "generate hundreds of thousands of tokens in a single trajectory" and solve hard problems in one go.

For an agent built on Argon, that changes the operating limits you plan around:

  • Runtimes - one response can now run far longer than before, so raise client timeouts and stream output instead of waiting for a single payload.
  • Cost caps - a maximum-length response is billed for a full 1M output tokens, so set a maximum output token value on every call rather than relying on the model to stop early.
  • Cached context - cached input tokens are priced at 95% off the input price, which favors agents that resend the same long context on every turn.
  • Grading - long outputs such as whole-library migrations need graders that check the complete result, like a full diff and a build, rather than spot checks of a few sections.

How Strong Is Gemini 4 Argon at Cybersecurity?

Google trained Argon for cybersecurity defense: it can autonomously find, validate, and patch critical software vulnerabilities, and it ties GPT-6 Astra for first place on CWE-bench v1, which measures vulnerability remediation.

For trusted defenders and its own teams, Google is releasing Argon without cyber guardrails. Wiz is already using it in its Scan for Good initiative, where the model found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide, a risk earlier frontier models had missed.

Access runs through the Fairwind Program, which gives governments, critical infrastructure operators, and core technology platforms early access to frontier models for defense. Google DeepMind says it works with over 650 partners, who can also use Argon inside CodeMender, its code security agent.

  • Dual-use tasks only - partners may run authorized threat simulation, reverse engineering, and malware analysis for defensive and academic research.
  • Security controls - partners agree to user-level authentication and phishing-resistant MFA, and may grant Argon access only to internal cybersecurity, incident response, or penetration testing teams.
  • No resale - partner organizations cannot share, redistribute, or sell access to the model.

What Safeguards Ship With Gemini 4 Argon?

Google lists the safeguards it is strengthening before broad release:

  • Misuse - Argon is designed to refuse cyber and CBRN attack requests while preserving legitimate dual-use research, and Google monitors the model's internal activations to spot misuse.
  • Prompt injection - Google calls Argon its most resilient model yet against indirect prompt injection, leading Gray Swan's Indirect Prompt Injection benchmark after automated red teaming and adversarial training.
  • Misalignment - monitors watch Argon's chain of thought and actions and stop execution when the model steps beyond the user's intentions.
  • Sandboxes - Google is isolating and sealing its sandboxed environments before high-risk training or evaluations begin.

Those controls protect Google's own deployment. Your agent still needs its own checks: prompt injection testing covers the attacks, and the OpenAI Hugging Face incident shows why a record an agent can write is weak evidence of what it did.

Note

Note: Google's monitors watch what Argon does inside Google's systems. For agents you build that call tools and change files, TestMu AI Agent Assurance grades each run against what actually changed. Try TestMu AI free!

When Is Gemini 4 Argon Available, and What Does It Cost?

Google has not announced a public release date. It says it is taking part in the U.S. government's voluntary process for pre-release model access while it widens access in this order:

  • Trusted cyber defenders in the Fairwind Program, plus trusted testers.
  • Paid API customers and Google AI Ultra subscribers.
  • Developers, enterprises, and consumers more broadly.

Google's announcement sets the API prices:

  • Introductory price - $2 per million input tokens and $10 per million output tokens.
  • After the introductory period - $4 per million input tokens and $20 per million output tokens.
  • Cached input - 95% off the input token price.

How Should You Test Agents Built on Gemini 4 Argon?

Argon leads Google's table on DeepSWE v1.1, and its score there still leaves roughly one long-horizon coding task in five unsolved. Code an Argon agent writes needs checking before it ships, the same as code from any other model.

  • Benchmark on your own tasks - run your evaluation suite on Argon and on your current model, weighting the terminal and computer-use tasks where Google's table shows Argon trailing.
  • Cap long outputs - set maximum output tokens and timeouts per call, and test how your app streams a response of several hundred thousand tokens.
  • Rerun prompt-injection tests - Google reports leading robustness on Gray Swan's benchmark, but your agent's tools and data differ, so attack your own setup.
  • Check refusals on security tasks - the broad release keeps the cyber guardrails that Fairwind access removes, so test vulnerability-research prompts against the version you will actually get.
  • Verify what coding agents ship - check the running app, not only the diff and the unit tests.

Kane CLI covers that last check. It validates rendered UI in a real Chrome browser from a plain-English objective, runs from your terminal, a coding agent's loop, or CI, and returns a pass or fail backed by an evidence pack. The same pattern with Google's coding agent is shown in Gemini CLI with Kane CLI.

For chat, voice, and phone agents moving to a new model, Agent Testing scores conversations on 9 quality metrics, including hallucination, and its testing agents include a Security Researcher persona that probes for prompt injection and data exfiltration. The Agent Testing CLI runs those evaluations from a terminal or a CI pipeline.

Kane CLI - Testing Agent in Your Terminal

Getting Started With Gemini 4 Argon

Until Argon reaches paid API customers, build the comparison now: collect the tasks your agent fails today, run them on your current model, and keep the results ready to rerun on Argon when access opens. Keep the UI checks in CI with Kane CLI remote execution so every model switch is tested the same way.

Author

...

Chaitanya Sharma

Blogs: 18

  • Linkedin

Chaitanya Sharma is an AI Product Manager at TestMu AI (formerly LambdaTest), where he builds agentic AI capabilities focused on computer vision and multi-modality, moving testing beyond static script execution toward autonomous, agent-driven workflows. Before TestMu AI he shipped 135+ features at Sprinklr for a no-code community and website builder used by Fortune 500 enterprises including Dell, Samsung, and Polestar. At Policybazaar he led the zero-to-one launch of a digital lending and insurance marketplace embedded in Bahrain's dominant payments app, building a risk-intelligence engine that compressed loan-approval times by 80%. He explored machine learning and NLP through research at the University of Cambridge, and holds a B.Tech from Delhi Technological University.

Reviewer

...

Sirajuddin Khan

Reviewer

  • Linkedin

Sirajuddin Khan is Vice President of Product Management at TestMu AI (formerly LambdaTest), where he drives the company's agentic AI product strategy, building a suite of autonomous agents that includes Agentic Browsers and Agentic Visual Testing and shifting the unit of work from test execution to autonomous outcomes. One of the company's earliest product leaders, he has owned the roadmap for the high-performance execution cloud and grew the cross-browser testing products from early adoption to market leadership. He brings over a decade of experience across SaaS, B2B, and eCommerce, with earlier product roles at Wydr and ShopClues, where his catalog and search work cut delivery SLAs and lifted seller activity. Sirajuddin holds an MBA in Information Technology from Sikkim Manipal University and a B.Tech in Computer Science Engineering from Maharshi Dayanand University.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Gemini 4 Argon FAQs

Did you find this page helpful?

More Related Blogs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests