1 October 2026

Gemini 4 Argon resets the model shortlist, not the production stack

Google’s strongest model arrives with aggressive pricing, long outputs and notable benchmark results, but restricted access makes this a reason to prepare an evaluation path rather than migrate.

Virtual Arc · Editorial image

What launched

Google announced Gemini 4 Argon on September 30 for long-running software engineering, enterprise knowledge and cybersecurity work. It will launch at $2 per million input tokens and $10 per million output tokens, with cached input discounted by 95%; standard pricing later rises to $4 and $20. Its output limit reaches one million tokens, but access currently remains restricted to Fairwind participants.

Evidence is mixed

Google reports 77.9% on DeepSWE v1.1. Vals independently placed Argon first on its broad work index, but only fifth on Terminal-Bench 4.0 and near the bottom on its computer-use evaluation. For operators, that is a reminder that a leading aggregate score does not establish end-to-end reliability on a specific workflow.

What we would do

We would prepare an adapter, representative task set and evaluation budget immediately. We would keep the existing production route until public API access lets us measure latency, tool stability, retry rates and actual cost per accepted result. Argon belongs in the next bake-off, but not yet in a delivery commitment.

Our take

Gemini 4 Argon is the first launch in this cycle that would make us reopen a frontier-model bake-off, but it would not make us move production traffic. Google is advertising introductory pricing of $2 per million input tokens and $10 per million output tokens, a 95% cached-input discount and a one-million-token output ceiling; independent Vals results also place Argon first on its broad work index. That combination could materially lower cost per completed long-horizon task if it survives contact with our workloads. But the model is currently restricted to trusted cyber defenders, has no public API model ID, and the independent picture is mixed: strong knowledge-work and code-migration results, but weaker placement on terminal coding and computer-use tests. Virtual Arc would build the adapter and frozen evaluation suite now, reserve capacity when paid API access opens, and run parallel trials on copies of production requests. We would not migrate a default route, promise an Argon-backed feature or rewrite orchestration around million-token trajectories until we have measured p95 latency, retries, tool-call failure rates and total cost per accepted result. The right move is to prepare, not to commit to a model we still cannot reliably buy and operate.

Sources
  1. Gemini 4 Argon: our next era of frontier intelligence
  2. Gemini 4 Argon model evaluation
  3. Gemini 4 Argon benchmarks, cost and capabilities
  4. Google announces Gemini 4 Argon AI model, but you cannot use it yet

← All posts