[models] · · 2 min read
Google drops three new Gemini models — but no 3.5 Pro yet
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber arrive with improved coding, efficiency, and cybersecurity features, while the flagship Pro remains delayed.
By ByteBulletin Editors · Editorial Team
Google DeepMind released three new Gemini models on Tuesday, focusing on efficiency, cost, and specialized security: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The lineup notably lacks the long-awaited Gemini 3.5 Pro, which has been delayed internally.
What's new
Gemini 3.6 Flash is positioned as Google's "workhorse model," bringing enhanced coding, knowledge work, and multimodal capabilities while cutting token usage by up to 17%, making it cheaper than the 3.5 Flash. This is the model most developers building production applications will likely reach for — faster responses, lower cost, and improved code generation.
Gemini 3.5 Flash-Lite is the most cost-effective option in the Flash family, optimized for high-volume, latency-sensitive tasks where budget is paramount.
Gemini 3.5 Flash Cyber is a specialized model fine-tuned for finding and fixing cybersecurity vulnerabilities. It's priced competitively but will be available only to governments and trusted partners in a limited pilot program — a move that signals Google's push into regulated, high-stakes AI use cases.
The missing flagship
The elephant in the room is Gemini 3.5 Pro. Google teased it alongside the 3.5 Flash release in May, saying it would roll out "next month." Bloomberg later reported internal delays as the model struggled to meet performance benchmarks. Logan Kilpatrick, Google DeepMind's product lead, said the company is still testing 3.5 Pro with partners and hopes to "land soon."
Meanwhile, competitors have been shipping fast: OpenAI released GPT-5.5 and is rolling out GPT-5.6, while Anthropic launched Claude Opus 4.8 and Claude Sonnet 5. The pressure on Google to deliver a flagship that matches the frontier is mounting.
What developers should care about
For developers building agentic workflows or production applications, the new Flash models offer tangible improvements: lower cost per token, better coding performance, and a specialized cybersecurity model that could help automate vulnerability remediation. The missing Pro model, however, means developers needing top-tier reasoning still have to look elsewhere — or wait.
Kilpatrick also noted that Google has begun its "most ambitious pre-training run yet" for Gemini 4, suggesting the next generation is already in the works.
SHARE
RELATED

[models] ·
Escha Labs ships 2-bit Qwen3.6-35B-A3B build that runs on a 24GB GPU
A 12.3GB MoE quant packs a 35B model onto consumer cards, with quality within noise of FP8 on most benchmarks.

[models] ·
ByteDance Is Training a 10 Trillion-Parameter Model, Leaning Into Scale to Chase Anthropic
The TikTok parent is reportedly pre-training a massive model that could rival Anthropic's Mythos 5, signaling a new phase in the global AI race.

[models] ·
Alibaba’s Qwen3.8-Max goes open-weight, claiming Claude-rivaling performance
Alibaba releases its largest open-weight model yet, Qwen3.8-Max, claiming it rivals Anthropic’s Claude Fable 5, with weights due next week.
