Google’s answer to a summer of being outrun is not a bigger model. It is cheaper ones.
On Tuesday the company launched three new Gemini models, all at its fast, low-cost “Flash” tier, a day before Alphabet reports earnings. The message is efficiency over raw power. The catch is what is still missing.
Cheap, fast, and everywhere but the top
The workhorse is Gemini 3.6 Flash. Google says it does better coding and knowledge work than its predecessor while using about 17% fewer output tokens, and it costs less. That is $7.50 per million output tokens, down from $9. Its knowledge now runs to March 2026.
Alongside it sits 3.5 Flash-Lite, the fastest of the family at 350 tokens a second, and cheaper still. Both lean into one bet: that most AI work does not need a frontier brain, just a quick, affordable one.
That bet has a champion. “Companies are already blowing through their annual token budgets, and it’s only May,” chief executive Sundar Pichai said earlier this year. A mix of Flash models, he told Business Insider, could save firms over $1 billion a year.
A cheaper shot at Mythos
The most pointed release is Gemini 3.5 Flash Cyber. It is tuned to find and patch software vulnerabilities, running inside Google’s CodeMender agent. Google calls it a “cost-efficient” alternative to large security models.
The unnamed target is Anthropic’s Mythos, which costs $10 per million input tokens and $50 per million output, as The Verge notes. Google claims Flash Cyber matches frontier performance on a key benchmark at a fraction of the price, a direct swipe at Anthropic’s lead in AI-driven security.
Because a bug-finder also helps attackers, Google is keeping it on a leash. Flash Cyber goes only to governments and trusted partners, in a limited pilot.
The model they did not ship
Then there is the absence. Gemini 3.5 Pro, the flagship Google promised for June, is still in testing, reportedly held back after falling short on coding. Google has no model in the public top ten.
The timing stings. In roughly a week, xAI’s Grok 4.5, three versions of OpenAI’s GPT-5.6, and Moonshot’s Kimi K3 all shipped. Anthropic’s Fable 5, meanwhile, sits atop the leaderboards, as Reuters reported.
Gemini 4, on paper
Google’s reply is to point further ahead. It says it has begun its “most ambitious pre-training run yet” for Gemini 4. That is a statement of intent, not a shipped capability.
The efficiency drive runs deeper than software. Google is also building a custom chip to serve Gemini far more cheaply. One caveat is worth keeping, though: every benchmark here is Google’s own, and no outsider has checked them yet.
The bet
The plan is coherent. In a year when companies are counting tokens, cheap and fast can win the middle of the market while the flagship catches up. Whether it holds rests on two things: 3.5 Pro finally shipping, and Gemini 4 turning out to be more than a training run.
Get the TNW newsletter
Get the most important tech news in your inbox each week.