
Google has released three new Gemini models in a clear attempt to shift the AI conversation from raw frontier bragging to speed, cost and security. The new lineup includes Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber.
In its official announcement, Google said Gemini 3.6 Flash improves coding and knowledge-work performance while using fewer tokens than 3.5 Flash. Google says the model consumes 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index and is priced at $1.50 per million input tokens and $7.50 per million output tokens.
Gemini 3.5 Flash-Lite is aimed at high-throughput workloads. Google says it can reach 350 output tokens per second, with pricing at $0.30 per million input tokens and $2.50 per million output tokens. The third model, Gemini 3.5 Flash Cyber, is a limited-access cybersecurity model designed to find, validate and patch vulnerabilities.
The timing is not accidental. Developers are increasingly choosing models based on cost per task, not only benchmark rank. Chinese models such as Kimi K3 and Qwen variants are putting pressure on U.S. labs by offering strong performance at lower prices. Google is responding with models that promise lower token usage and better throughput.
That is why Gemini 3.6 Flash matters even if it is not a new flagship model. For many businesses, the winning model is not always the smartest one in a benchmark table. It is the one that can run thousands or millions of tasks reliably without making the bill explode.
This fits the same infrastructure reality behind our Kimi K3 capacity story. Cheap AI creates demand, and demand turns efficiency into a competitive weapon.
Gemini 3.5 Flash Cyber may be the most strategically interesting part of the announcement. Google says it is fine-tuned for security work and will be offered first to governments and trusted partners through a limited-access pilot because of the dual-use nature of cyber capability.
Google is positioning Flash Cyber as a cost-efficient alternative to heavier cyber models such as Anthropic’s Mythos. The model is meant to work with Google’s CodeMender platform, scanning code, finding vulnerabilities and helping patch them.
This is arriving at exactly the moment AI cyber capability is under intense scrutiny. OpenAI’s Hugging Face incident shows that models can chain vulnerabilities and act unexpectedly. Google is trying to show the defensive side of that same capability: AI models that help security teams find and fix problems faster.
The weaker part of the announcement is what Google did not release. Gemini 3.5 Pro is still not generally available, and that absence will attract attention because OpenAI, Anthropic, xAI, Moonshot and other labs have kept the pressure high.
Google says it has started its most ambitious pre-training run yet for Gemini 4. That may reassure some developers that a larger jump is coming, but it also makes this release feel like a bridge: useful, practical and cost-focused, but not the full comeback some expected.
Still, Google has an advantage many AI labs do not: distribution. Gemini can reach Search, Android, Workspace, Vertex AI, Google Cloud and consumer apps. If the models are fast and cheap enough, distribution can matter as much as leaderboard position.
What Developers Should Watch
The first thing to watch is real task cost. Google is talking about token efficiency, fewer tool calls and lower prices. Developers will quickly test whether those savings appear in actual coding agents, document workflows, customer-support systems and data pipelines.
The second thing is reliability. A cheaper model that fails more often can become more expensive after retries, human correction and broken workflows. The third is whether Flash Cyber proves useful enough for governments and trusted partners to trust it with real vulnerability work.
Google is clearly trying to make Gemini more practical and deployable. In this phase of AI, that may be exactly the right fight. Not every customer needs the biggest model. Many need a model that is fast, good enough, safer to deploy and cheaper to run all day.