Gemini 4 Argon: Google AI finds and patches critical flaws

Gemini 4 Argon: Google AI finds and patches critical flaws

Google has unveiled Gemini 4 Argon, a new frontier AI model that the company says can find, validate and fix critical software vulnerabilities without human involvement. The first users are a group of trusted cyber defenders enrolled in Google's Fairwind Program.

Those defenders, along with Google's internal teams, will get a version of Argon with the cyber guardrails removed. Wider access comes later. Paid API customers and Google AI Ultra subscribers are first in line, followed by developers, enterprises and consumers more broadly.

Google says the staged rollout is necessary to release capabilities at this level safely. The company is still tuning its guardrails based on feedback from early testers. It is also participating in the U.S. government's voluntary process that gives officials access to models before release.

Safeguards before the wider launch

Ahead of general availability, Google says it is hardening protections against cyber misuse and against chemical, biological, radiological and nuclear (CBRN) abuse. Internal and external red teams, groups tasked with attacking a system to uncover its weak points, have tested these safeguards. Google adds that monitoring systems track Argon's reasoning and actions and can halt execution when necessary.

From healthcare bugs to Rust rewrites

One of the biggest technical changes is the output limit, the maximum amount of text the model can produce in a single response. It jumps from 64,000 tokens to 1 million. According to Google, this lets Argon reason and write across hundreds of thousands of tokens in one run.

Cloud security firm Wiz is already using the model in its Scan for Good program, which looks for high-risk exposures in critical public infrastructure and fixes them for free. Google says Argon uncovered a critical flaw in healthcare software used by hospitals around the world. The bug exposed sensitive personal information and had been missed by earlier frontier models. Google did not name the product or say whether a fix is available.

Inside Google, Argon agents rolled out memory optimizations across the company's data centers, freeing more than 300 TiB of memory. They also replaced 32,000 lines of SIMD code in the libgav1 video decoder with Rust, a language designed to prevent memory errors. The new decoder runs 2.7 times faster than the previous Rust port and produces identical video output.

Other agents are converting C and C++ code to Rust. This includes more than 800,000 lines for the Fuchsia Zircon kernel. None of these rewrites are in production yet. They are still going through audits, emulation testing and review.

Benchmark results

Google reports the following scores:

  • 77.9% on DeepSWE v1.1, a long-form software engineering test, which Google calls state of the art
  • 51.3% on Zapier's AutomationBench, ranking first
  • 91.7% on LVBench, a long-video benchmark, also claimed as state of the art
  • 68% on CWE-bench v1, which measures the ability to fix security vulnerabilities, tying for first place

The vulnerability discovery figures come from internal testing. Google's own evaluation covers complex codebases in 20 programming languages. In a black-box penetration test run by Wiz, where the model attacks live web systems without seeing source code, Argon outperformed 3.8 Flash Cyber at mapping the attack surface, finding flaws and producing proof-of-concept evidence.

Pricing

Argon launches at an introductory rate of $2 per million input tokens and $10 per million output tokens. Cached input tokens are 95% cheaper than standard input. Once the introductory period ends, prices double to $4 and $20.

Our take

Argon fits a pattern we have been tracking for months: AI is compressing the time between a bug existing and someone finding it. We recently reported that vulnerability disclosures have doubled as AI accelerates exploitation, and a model that can locate and patch flaws on its own could push that number higher still.

The decision to hand a guardrail-free version to selected defenders suggests Google is betting that giving the good guys a head start outweighs the risk. That approach depends heavily on voluntary commitments, much like the broader push for industry self-policing in the U.S.

It is worth watching whether the unnamed healthcare flaw gets disclosed and fixed, whether independent testing confirms Google's internal results, and how long the guardrails hold once Argon reaches paying customers.