Google announced Gemini 4 Argon — the first frontier model of the Gemini 4 generation — rolling out first to trusted cyber defenders through the Fairwind Program, with cyber guardrails removed for that cohort and for Google’s internal teams. The builder headline is clear even before public API access: 1M output tokens (up from 64K), intro pricing at $2 / $10 per million input/output tokens with 95% off cached input ($0.10), and a phased path that starts with Fairwind + US government voluntary pre-release review before paid API customers and Google AI Ultra.
Argon is not publicly available yet. There is no public date. Treat every peer score below as mostly vendor-reported until independent harnesses settle — and note early Artificial Analysis / HN debate that Sol 6.1 may still win on cost per task.
Image credit: Google
Key points
- What shipped (access-limited): Gemini 4 Argon — first of the Gemini 4 frontier generation; complex workflows across software engineering, enterprise knowledge work (legal/finance), and cybersecurity defense.
- Who has it now: Trusted cyber defenders via Fairwind (guardrails removed for them + Google internal). Next: paid API customers + Google AI Ultra. Broader developers/enterprises/consumers “as soon as possible” — no date.
- Intro API pricing (per 1M tokens): Input $2 · Output $10 · Cached input $0.10 (95% off). After intro: $4 / $20.
- Output limit: 1M tokens (up from 64K) — headroom for long single-trajectory reasoning and large agent rewrites.
- Internal Google uses (vendor): quantum algo optimization (40% better than a published baseline in one example); Argon agents freed 300+ TiB memory (est. 500 TiB–1 PiB); C/C++→Rust migrations (re2, libgav1, Fuchsia Zircon up to 800K+ lines); libgav1 Rust SIMD rewrite → 2.7× faster than the prior Rust port.
- Vendor benches (self-reported): DeepSWE v1.1 77.9% SOTA · Vals Index leading · Vals Finance Agent v2 · Harvey Legal Agent Benchmark · AutomationBench #1 at 51.3% · LVBench long video 91.7% SOTA · CWE-bench v1 ties first at 68%.
- Cyber: Wiz Scan for Good found a critical healthcare vulnerability prior frontier models missed (Google’s claim). Fairwind + internal get full cyber-capable Argon without cyber guardrails.
- Safeguards before broad availability: misuse (cyber/CBRN), prompt injection (Gray Swan IPI lead per Google), misalignment monitoring of CoT/actions, hardened sandboxes.
- Independent / secondary: AA Intelligence Index places Argon High roughly on par with Opus 5.5 high / Sol 6.1 / Astra on some axes, with HN skepticism on Google benchmaxxing and notes that Sol 6.1 may be cheaper per task. Mixed tallies: Argon trails on FrontierSWE v2, Terminal-Bench 4.0, OSWorld-2.0 in some vendor-reported stacks.
- Caveats: Not public yet. Benchmarks mostly vendor-reported. 9to6AI has not hands-on tested. Re-measure when API access lands.
What shipped
Google positions Argon as a frontier partner for long-horizon coding, enterprise knowledge work, and defensive cybersecurity — released carefully because those same skills raise misuse risk.
| Field | Value |
|---|---|
| Model | Gemini 4 Argon (first Gemini 4 frontier) |
| Access now | Fairwind Program cyber defenders + Google internal |
| Guardrails | Cyber guardrails removed for Fairwind + internal; kept for general release |
| Intro pricing | $2 in / $0.10 cached in / $10 out per 1M |
| Post-intro pricing | $4 in / $20 out per 1M |
| Max output | 1M tokens (was 64K) |
| Next surfaces | Paid API customers, Google AI Ultra (no date) |
| Author | Koray Kavukcuoglu, SVP Google DeepMind / Chief AI Architect |
Official announcement: blog.google … /gemini-4-argon.
What changed vs prior Gemini and the market
Cyber-first, not consumer-first. Unlike a simultaneous AI Studio / Gemini app drop, Argon opens inside Fairwind for vetted defenders, with Google citing US government voluntary pre-release engagement and a phased rollout. Secondary reporting (Straits Times / Reuters wire) frames the restriction as safety-driven: withhold the most capable cyber-capable model from general public access while testers find and patch real flaws.
1M output tokens. The practical change for agents is trajectory length — hundreds of thousands of tokens in one go for migrations, long research drafts, or multi-step remediation — not a chat novelty.
Pricing that matches today’s mid/frontier promo band. Intro $2/$10 with $0.10 cache reads sits next to GPT-6.1 Sol’s list rates. Post-intro $4/$20 aligns with Opus-class stickers. The bill that matters will be tokens per successful task, not the sticker alone.
Vendor coding and enterprise tables. Google claims DeepSWE v1.1 SOTA at 77.9%, AutomationBench #1 at 51.3%, leading Vals Index / Finance Agent v2 / Harvey Legal Agent, and LVBench 91.7%. Cyber: CWE-bench v1 tie at 68%; Wiz Scan for Good healthcare vuln story. Treat rival rows in a Google table as directional.
Where independent / secondary tallies push back. Early Artificial Analysis and HN debate: Argon High can look on par with Opus 5.5 high / Sol 6.1 / Astra on intelligence-style indexes, but Sol may win on $/task; Google benchmaxxing skepticism remains loud. Mixed vendor-reported coding stacks (FrontierSWE v2, Terminal-Bench 4.0, OSWorld-2.0) show Argon trailing some peers — Google’s own launch materials do not claim a clean sweep across every agent harness.
Internal proof points (unverifiable from outside). Quantum subroutine spacetime optimization (+40% vs a published baseline in one example); fleet memory optimizations (≥300 TiB freed, est. 500 TiB–1 PiB); large C/C++→Rust agent migrations including Fuchsia Zircon (800K+ lines) under heavy audit; libgav1 SIMD rewrite in safe Rust at 2.7× the prior Rust port. Useful as intent signals, not as your eval.
What it means for builders
If you expected a public Gemini 4 API today. You do not have it. Plan around Fairwind exclusivity and watch for paid API + AI Ultra — Google has not pinned a calendar.
If you run defensive security programs. Fairwind is the gate. Google’s pitch is autonomous find/validate/patch plus Wiz Scan for Good early results. Apply only if you fit their trusted-defender criteria; do not assume AI Studio access.
If you are comparing Sol 6.1 / Opus 5.5 / Astra on price. Intro Argon matches Sol’s $2/$10 sticker and undercuts Astra’s $10/$50. Independent AA/HN commentary already argues Sol can be cheaper per task once token use is counted. Budget on post-intro $4/$20 until Google states how long the discount lasts.
If long agent trajectories are your bottleneck. 1M max output is the differentiator to test first when access lands — migrations, multi-file remediations, long legal/finance research loops. Measure wall-clock and output spend; a million-token answer at $10/M is a $10 line item at intro rates.
If you buy vendor benches at face value. Don’t. DeepSWE / AutomationBench / Vals / LVBench leads are Google-reported. AA and other harnesses are still settling. Re-run your own coding and cyber suites before flipping production defaults.
Safety / compliance teams. Expect misuse refusals (cyber/CBRN), prompt-injection hardening (Gray Swan IPI lead per Google), CoT/action misalignment monitors, and hardened sandboxes before broad availability. Fairwind’s no-cyber-guardrail mode is explicitly not the consumer posture.
What to watch
- Public/paid API and Google AI Ultra availability date (still unannounced).
- Independent Artificial Analysis and third-party harness settlement vs Opus 5.5, Sol 6.1, and Astra — especially $/task and agent success, not sticker price.
- Whether Fairwind + Wiz-style cyber wins generalize beyond Google’s named examples.
- How long introductory $2/$10 lasts before $4/$20.
- Whether 1M-output trajectories change real migration/remediation workflows once outsiders can measure them.



