Google has unveiled Gemini 4 Argon, the first new generation of its flagship AI model since Gemini 3 last November. A small group of cybersecurity defenders gets it first. Paid API customers and Google AI Ultra subscribers are next, Google said on Wednesday.
Koray Kavukcuoglu, who now runs Google DeepMind day to day, called Argon the company’s “next era of frontier intelligence” in the announcement. Google says the model is built for long, complex workflows in software engineering, legal and finance work, and cybersecurity defence.
Cyber defenders first
The first users come through Google’s Fairwind Program for “trusted cyber defenders”. Google is also taking part in the US government’s voluntary process for pre-release model access. It will widen access after more testing.
Those defenders get Argon without its cyber guardrails, so they can use its full abilities, Google said. Security firm Wiz is already using it to find flaws through its Scan for Good initiative. The model found a critical vulnerability in healthcare software used by hospitals worldwide, one that earlier frontier models had missed, according to the company.
Tulsee Doshi, head of Gemini products at Google DeepMind, spoke to Axios.
“Argon is a well-rounded model that has frontier capabilities across several domains.”
The numbers Google gave
Argon sets a new high of 77.9% on DeepSWE v1.1, a test of long software engineering tasks, according to Google. It ties for first place, at 68%, on CWE-bench, which tests how well models fix security flaws. It beat OpenAI’s GPT-6 Astra on several coding and knowledge-work benchmarks, Axios reported.
Google is also raising the model’s output limit to 1 million tokens, up from 64,000. Thousands of Google staff already use Argon, and teams of Argon agents freed more than 300 TiB of memory across Google’s data centres, the company said. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens.
Doubts inside Google
Some Google employees are sceptical. The model does well on benchmarks but less well when staff put it to work. It also struggles with some coding tasks, Bloomberg reported, citing people with direct access to the effort. Julia Love and Davey Alba also reported that Google dropped Gemini 3.5 Pro. It had promised that model for June.
Google told Bloomberg it would be inaccurate to say Gemini 4 underperforms in areas such as coding. One employee told Bloomberg there is “large consensus” inside the company that the model is at the frontier.
Kavukcuoglu said last week that Google wanted Gemini 4 out well before year-end. Rivals have been dealing with safety problems of their own. OpenAI cancelled its next model this week after it failed internal safety tests. It also faces a lawsuit over agents that hacked Hugging Face.
Get the TNW newsletter
Get the most important tech news in your inbox each week.