Google is calling all three of its latest Gemini models the most advanced work it’s ever released. That’s the same claim it makes with every launch. So I dug for the single figure that genuinely counts, and it turned out to be tucked inside a token-usage number that most readers will breeze right past.
![Google shipped three new Gemini models, but only one changes your bottom line [2026] 28 google dropped 3 new gemini models and only one earns the hype 2026 1 Google is calling all three of its latest Gemini models the most advanced work it's ever released. That's the same claim it makes with every launch. So I dug for the single figure that genuinely counts, and it turned out to be tucked inside a token-usage number that most readers will breeze right past.](https://egamers.io/wp-content/uploads/2026/07/google-dropped-3-new-gemini-models-and-only-one-earns-the-hype-2026-1.png)
The trio consists of Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber — three models built for three distinct jobs, yet only one of them shifts what running it at scale will cost you.
The 17% that pays your bill
Begin with 3.6 Flash, which Google positions as its workhorse. The sell is stronger coding, sharper knowledge work and improved multimodal handling versus the model it succeeds.
Now for the detail that deserves your focus. Per the Artificial Analysis Index, 3.6 Flash trims output token usage by 17% relative to 3.5 Flash. On certain benchmarks, such as DeepSWE by Datacurve, Google claims that figure climbs to as much as 65%.
![Google shipped three new Gemini models, but only one changes your bottom line [2026] 29 google dropped 3 new gemini models and only one earns the hype 2026 2 Google is calling all three of its latest Gemini models the most advanced work it's ever released. That's the same claim it makes with every launch. So I dug for the single figure that genuinely counts, and it turned out to be tucked inside a token-usage number that most readers will breeze right past.](https://egamers.io/wp-content/uploads/2026/07/google-dropped-3-new-gemini-models-and-only-one-earns-the-hype-2026-2.png)
Anyone who’s seen a Flash model chew through tokens narrating its own reasoning before it actually answers will understand why this matters. Fewer output tokens paired with a lower per-token cost isn’t a flashy feature — it’s the number that lands on your bill.
In Google’s own words: “3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance. According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash, and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token.”
![Google shipped three new Gemini models, but only one changes your bottom line [2026] 30 google dropped 3 new gemini models and only one earns the hype 2026 3 Google is calling all three of its latest Gemini models the most advanced work it's ever released. That's the same claim it makes with every launch. So I dug for the single figure that genuinely counts, and it turned out to be tucked inside a token-usage number that most readers will breeze right past.](https://egamers.io/wp-content/uploads/2026/07/google-dropped-3-new-gemini-models-and-only-one-earns-the-hype-2026-3.png)
Flash-Lite is built for one thing: speed
Next comes 3.5 Flash-Lite, tuned for low-latency tasks, chat and document processing. Google bills it as the quickest and most cost-effective option in its 3.5 family.
The headline figure here is 350 output tokens per second — once more sourced from the Artificial Analysis Index. Google adds that it surpasses earlier Flash-Lite generations in agentic workflows, and notes specifically that it “significantly” beats its predecessor at the thinking level.
This is the model you turn to when someone’s waiting on an answer and every second of delay carries a cost. It isn’t built to out-think the workhorse — it’s built to respond before you register the wait.
Google’s line: “3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second according to the Artificial Analysis Index, also significantly outperforming prior Flash-Lite generations in agentic workflows.”
The security one is a model plus an agent, not just a model
The third entry, 3.5 Flash Cyber, is the one likely to be misunderstood. It’s fine-tuned to help detect and patch cybersecurity vulnerabilities, and it does so at a lower per-token cost than the bigger models.
But Google is candid that the model on its own isn’t the whole picture. It couples a specialized cyber-focused model with its CodeMender code security agent, and the reasoning is explicit: security work requires the model orchestrated alongside agent infrastructure rather than deployed by itself.
Google’s description: “3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure. We’re introducing a combination of a new, highly efficient, specialized cyber-focused model paired with our CodeMender code security agent that delivers competitive performance at the frontier.”
Look closely at the wording. “Competitive performance at the frontier” — not best in the world. Whenever a company softens its own boast, take the hedge at face value.
Which one to actually use
Running a coding assistant or anything multimodal at scale? Put 3.6 Flash at the top of your test list, and that 17% token reduction is precisely why. If your application succeeds or fails on response speed, Flash-Lite and its 350 tokens per second is your choice. And for vulnerability work, resist the urge to judge Flash Cyber on its own, since Google never intended it to operate that way.
Three models: one that rewrites the cost equation and two that keep to their lanes. Google says you can find out more on its announcement page.
Source: GSMarena.




STAY ALWAYS UP TO DATE