top of page

Google launches Gemini 4 Argon with frontier performance in coding, finance, legal work and cybersecurity

28 minutes ago
4 min read
Google Gemini 4 Argon frontier model for coding, finance, legal work and cybersecurity

Google has launched Gemini 4 Argon, a new frontier model built for complex, long-running professional workflows across software engineering, finance, legal work and cybersecurity.


Argon raises the maximum output limit to 1 million tokens, up from 64,000 tokens in Google's previous configuration. The expansion applies to generated output rather than simply increasing input context, giving the model substantially more room for extended reasoning, coding, checking and revision within one task.


Google is initially restricting access to trusted cybersecurity defenders through its Fairwind Program while additional safety testing continues. Wider availability is planned for developers, enterprises and consumers, beginning with paid API customers and Google AI Ultra subscribers.


Introductory API pricing starts at $2 per million input tokens and $10 per million output tokens. Google says those rates will later double to $4 and $20 respectively, while cached input receives a 95% discount during the introductory period.


··········


GEMINI 4 ARGON AT A GLANCE


........


Metric

Gemini 4 Argon

Model class

Frontier model

Introductory input price

$2 / 1M tokens

Introductory output price

$10 / 1M tokens

Maximum output

1M tokens

DeepSWE v1.1

77.9%

Vals Index

68.9%

Finance Agent v2

65.4%

AutomationBench

51.3%

LVBench

91.7%

CWE-bench v1

68.0%


........


The 95% cache discount turns the introductory $2 input rate into an effective $0.10 per million cached input tokens. Google has not specified how long the introductory pricing period will last.


··········


THE 1 MILLION TOKEN OUTPUT LIMIT EXPANDS LONG-HORIZON EXECUTION


Large context windows determine how much information a model can receive; output limits determine how much work it can generate before a single trajectory ends. Argon's move from 64K to 1M output tokens gives software-engineering and research agents substantially more room to inspect, plan, edit, test, diagnose failures and continue working without repeatedly restarting the task through separate model calls.


The specification is useful only when a workload genuinely requires extended execution. Very long generations increase latency and inference cost, so the practical advantage is the removal of a hard ceiling rather than an assumption that every task should consume hundreds of thousands of tokens.


··········


ARGON'S STRONGEST RESULTS CLUSTER AROUND PROFESSIONAL AND AGENTIC WORK


Google reports 77.9% on DeepSWE v1.1 for long-horizon software engineering, 68.9% on the Vals Index across professional knowledge work and 65.4% on Finance Agent v2 for multi-step financial research. The model also reaches 51.3% on AutomationBench and 91.7% on LVBench.


The profile is not a clean sweep. Argon scores 57.4% on Terminal-Bench 4.0, below Claude Opus 5.5's 66.4% in Google's published comparison, while its 55.0% on FrontierSWE v2 trails both GPT-6 Astra and Opus 5.5 in the same table. The benchmark mix therefore points to workload-specific strengths rather than a universal lead across every coding environment.


··········


GOOGLE IS ALREADY USING ARGON FOR LARGE INTERNAL ENGINEERING TASKS


Google says Argon agents analyzed profiling telemetry across its infrastructure and identified memory optimizations that freed more than 300 TiB of memory after deployment, with estimated potential savings between 500 TiB and 1 PiB.


The company is also using Argon on C/C++-to-Rust migrations extending to more than 800,000 lines of code in the Fuchsia Zircon kernel. In Google's libgav1 work, Argon replaced roughly 32,000 lines of SIMD code and the resulting memory-safe implementation ran 2.7× faster than the previous Rust port while producing identical video output.


These are company-reported internal results rather than standardized independent benchmarks, but they show the scale of task Google is targeting: long-running engineering workflows where an agent must continue working across many intermediate operations before reaching a finished result.


··········


CYBERSECURITY CAPABILITY IS DRIVING THE LIMITED INITIAL ROLLOUT


Google says Argon can autonomously identify, validate and patch software vulnerabilities and reaches 68% on CWE-bench v1. Trusted defenders in the initial Fairwind rollout receive access to the model's full defensive capabilities while broader deployment remains more restricted.


Google also describes Argon as its most resilient model so far against indirect prompt injection. That risk becomes more consequential as agents read external documents, websites and tool outputs, because malicious instructions embedded in those sources can attempt to redirect the model away from the user's intended objective.


The rollout therefore combines model-level adversarial training with monitoring systems intended to detect behavior that moves outside the requested objective and terminate execution when necessary.


··········


THE INTRODUCTORY PRICE WILL DOUBLE


Argon's launch pricing is $2 per million input tokens and $10 per million output tokens. After the introductory period, Google says the rates will rise to $4 and $20 respectively.


Data Studios illustrative calculation:

10M uncached input tokens + 2M output tokens at introductory pricing = $40


The same token volume at post-introductory pricing = $80


The calculation excludes caching and does not assume that different models consume identical token volumes to complete the same task. With long-horizon agents, cost per completed task remains more informative than nominal token price alone.


··········


ARGON EXTENDS THE FRONTIER RACE FROM CHAT QUALITY TO COMPLETE WORKFLOWS


Gemini 4 Argon is designed around a production constraint that becomes more important as AI systems move beyond short-form interaction: how much complex work a model can reliably carry from beginning to end.


Its 1M-token output limit, strong results in finance and professional-work agents, long-horizon coding performance and cybersecurity specialization give Google a model optimized for extended tool-using execution rather than simply better single-turn answers. Argon does not lead every coding benchmark, broad access is not yet available and its announced API price will double after the introductory period, so the deployment case remains workload-specific.


For developers and enterprises, the practical change is a higher ceiling for software, research, document and security workflows that require the model to continue operating across many intermediate steps before delivering a finished result.


··········


FOLLOW US FOR MORE.


DATA STUDIOS


datastudios.org

bottom of page