Google Ships Three Gemini Flash Models, Holds Back 3.5 Pro

Google Ships Three Gemini Flash Models, Holds Back 3.5 Pro

The new 3.6 Flash cuts output token use by 17% and undercuts its predecessor on price, while the flagship Pro update slips amid internal delays.

Richard Miniter
First Published: July 22, 2026, 6:37 AM ETUpdates (1): July 22, 2026, 6:37 AM ET

Google DeepMind released three new Gemini models built for cost and speed, while withholding the flagship 3.5 Pro update its rivals have raced to match.

The lineup pairs a cheaper, more efficient workhorse with a low-latency lite model and a specialized cybersecurity system, the company said in a blog post on July 21, 2026 at 11:16 AM ET. Gemini 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, priced at $1.50 per million input tokens and $7.50 per million output tokens.

“Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance,” wrote Tulsee Doshi, senior director of product management on the Gemini team, framing the Flash series as the balance point between cost and quality for agentic workflows.

Gemini 3.5 Flash-Lite anchors the low-cost end of the range. Google said the model runs at 350 output tokens per second, the fastest in the 3.5 series, and lists at $0.30 per million input tokens and $2.50 per million output tokens. It aims at high-throughput jobs like agentic search and document processing.

The third model targets security work. Gemini 3.5 Flash Cyber, tuned to find and fix code vulnerabilities and paired with Google’s CodeMender agent, will reach only governments and trusted partners through a limited-access pilot, according to techcrunch. Google positioned the combination as competitive at the frontier of automated code security.

The launch drew attention as much for the omission as the releases. The update left out the long-awaited refresh of Gemini Pro, Google’s highest-capability line for complex reasoning, which last shipped in February 2026, according to techcrunch. Bloomberg reported last week that Google had hit internal delays as the 3.5 Pro struggled to meet performance targets, according to techcrunch.

The gap matters because rivals have not paused. OpenAI has released GPT-5.5 and started rolling out GPT-5.6, while Anthropic launched Claude Opus 4.8 and Claude Sonnet 5 and widened access to its Fable 5 model, according to techcrunch. Google teased Pro during its May 2026 Flash release, saying the version was already in internal use and would roll out the following month.

For developers who build on these models, the shift is practical rather than abstract. Lower token consumption and reduced per-task cost make automated agents cheaper to run at volume, and the Flash tier now carries computer use as a built-in client-side tool through the Gemini API and Gemini Enterprise.

Google has staggered its Gemini releases this way before, leading with cost-efficient Flash tiers ahead of heavier Pro models. On benchmarks cited in the announcement, 3.6 Flash scored 49% on DeepSWE against 37% for 3.5 Flash and 83.0% on OSWorld-Verified against 78.4%, gains the company tied to fewer reasoning steps and tool calls.

The next decision rests with Google on when 3.5 Pro reaches the public. DeepMind product lead Logan Kilpatrick said on July 21, 2026 that the company is testing 3.5 Pro with partners and hopes to “land soon,” without naming a firm date, according to techcrunch. He added that the team has begun its most ambitious pre-training run yet, for Gemini 4.

The race to name the smartest machine is older than the current model wars. In 1965, computer scientist Herbert Simon predicted that machines would do any work a human could do within 20 years, a forecast that passed without arriving. Decades of AI research swung between booms and funding winters known as “AI winters,” the first striking in the 1970s after early promises outran results, according to britannica.com. The pattern rewarded labs that shipped steady, usable tools over those that chased the grand claim.


Research