Google released three lightweight AI models Tuesday but offered no update on when its delayed flagship Gemini 3.5 Pro will arrive.
Google released three lightweight AI models Tuesday but offered no update on when its delayed flagship Gemini 3.5 Pro will arrive.

Google released three lightweight AI models Tuesday but offered no update on when its delayed flagship Gemini 3.5 Pro will arrive.
Google's new Gemini 3.6 Flash uses 17% fewer output tokens than its predecessor while undercutting comparable models from OpenAI and Anthropic on price, intensifying the AI pricing war as the company's flagship model remains delayed.
"We are focused on delivering the best price-performance ratio in the industry," a Google DeepMind spokesperson said.
Gemini 3.6 Flash, priced at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens — down from $9 for 3.5 Flash — scored 49% on the DeepSWE coding benchmark, up from 37% for its predecessor, and 63.9% on MLE Bench for machine learning research, compared with 49.7% previously. The model's knowledge cutoff advanced to March 2026 from January 2025. Google also introduced Gemini 3.5 Flash-Lite at $0.30 per 1 million input tokens and $2.50 per 1 million output tokens, which outperformed the prior-generation 3 Flash on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74% vs. 65.1%).
The launches come as Google's Gemini 3.5 Pro — originally slated for a June debut at the company's I/O developer conference — remains in partner testing with no public release date. Alphabet reports second-quarter earnings Wednesday, and Chief Executive Officer Sundar Pichai is expected to face questions about the delay. Google has already begun pre-training for Gemini 4, which the company described as its "most ambitious pre-training run yet."
The new lineup includes Gemini 3.5 Flash Cyber, a specialized variant for detecting and patching software vulnerabilities, available initially to governments and trusted partners through a limited-access pilot. Google's CodeMender tool uses multiple Flash Cyber agents to find and fix security issues at scale, the company said.
The pricing strategy reflects Google's bet that cost efficiency can offset its slower cadence in premium model releases. Artificial Analysis data shows Gemini Flash already undercuts comparable models from OpenAI, Anthropic and Chinese rivals on cost. Gemini 3.6 Flash is cheaper per task than OpenAI's GPT-5.6 Terra Max, Moonshot AI's Kimi K3 and Alibaba's Qwen 3.7 Max, according to the company.
The delay of Gemini 3.5 Pro has handed momentum to rivals. Anthropic's Mythos 5 and Fable 5 models were deemed so powerful by the US government as to be considered a national security risk, while OpenAI launched GPT-5.6 earlier this month after a similar government-requested security review. Chinese competitors are also gaining traction — Moonshot AI's Kimi K3 drew enough demand that the company limited new subscriptions because of capacity constraints, and Alibaba is teasing Qwen 3.8 Max, which it said trails only Anthropic's Fable 5 in overall performance.
For investors, the key question is whether Google's volume-driven pricing strategy can compensate for the absence of a top-tier model. Alphabet trades at roughly 22 times forward earnings, and Pichai has said companies could save upward of $1 billion per year by shifting most of their AI workloads to Gemini models. The company is also reportedly developing a specialized chip designed to run Gemini up to 10 times more efficiently, part of a broader push to lower AI serving costs.
This article is for informational purposes only and does not constitute investment advice.