Gemini 3.5 Flash is available since May 19, 2026. Launched at Google I/O, this model surpasses the former flagship Gemini 3.1 Pro on benchmarks while remaining four times faster.
When the Fast Model Beats the Premium Model
In the world of large language models, names matter. "Flash" variants have long designated lightweight models, designed for speed at the expense of quality. Gemini 3.5 Flash overturns this logic: Google claims performance that rivals large flagship models across several dimensions, while retaining the execution speed characteristic of the Flash range.
In its official blog post published on May 19, 2026, titled "Gemini 3.5: frontier intelligence with action," Google presents this new model as a major turning point in the development of more capable and intelligent AI agents. The wording is not insignificant: it's no longer just about answering questions, but about autonomously executing complex tasks.
Benchmarks Placing the Model in the Competition
On evaluations that matter for agentic use cases, Gemini 3.5 Flash outperforms Gemini 3.1 Pro on several key indicators: Terminal-Bench 2.1 (76.2%), GDPval-AA (1,656 Elo points), and MCP Atlas (83.6%). For multimodal understanding, it achieves 84.2% on CharXiv Reasoning.
Other tests shared by Google reveal 90.4% on GPQA Diamond, a benchmark that evaluates doctoral-level scientific reasoning, as well as 81.2% on MMMU-Pro for multimodal comprehension tasks. In terms of software development, Gemini 3.5 Flash reaches 78% on SWE-bench Verified.
However, these figures should be contextualized. As of May 2026, GPT-5.5 leads Gemini on several agentic benchmarks, and Claude Opus 4.7 maintains an advantage on certain programming tests, particularly SWE-bench Verified, with the lowest hallucination rate of the three. Gemini 3.5 Flash's advantage therefore clearly lies in its speed/price ratio at near-flagship quality, not in absolute market dominance.
Speed as a Central Argument
Gemini 3.5 Flash produces output tokens four times faster than other reference frontier models, making it the high-cadence engine necessary for real-world agentic workflows. In practice, this performance translates into the ability to handle tasks that used to take a developer days or an auditor weeks in a fraction of the time.
Google DeepMind emphasizes that the model achieves this level of performance through specific co-design between the model and the hardware infrastructure, enabling deeper reasoning capabilities to be trained more efficiently with each new generation. Economically, developer pricing is set at $1.50 per million input tokens and $9.00 per million output tokens via the API, about 40% less than Gemini 3.1 Pro.
Immediate Availability Across All Channels
Starting May 19, 2026, Gemini 3.5 Flash is accessible to billions of users worldwide: via the Gemini app and Google Search's AI mode for the general public, in Google Antigravity and the Gemini API for developers, and through the Gemini Enterprise Agent Platform for businesses.
For developers in particular, Google Antigravity 2.0, a standalone desktop application, now allows the deployment of collaborative agents capable of executing workflows and programming tasks simultaneously under human supervision. A "Managed Agents" API further simplifies deployment: a single API call is now sufficient to launch an agent capable of reasoning, using tools, and executing code in an isolated Linux environment.
Feedback from integrated partners on the official Google DeepMind page illustrates this potential. Box indicates that Gemini 3.5 Flash outperforms Gemini 3 Flash by 19.6% on its own internal evaluation dataset, and that its clients in life sciences are seeing a 96.4% improvement in scientific data extraction accuracy.
Gemini Spark, the First Beneficiary Agent
Gemini 3.5 Flash is not only deployed in developer tools. Google is using it to power Gemini Spark, a personal AI agent that operates 24/7, taking actions on behalf of the user in their daily digital life.
Spark is designed to automate recurring tasks, including analyzing bank card statements to detect hidden fees, monitoring email inboxes to flag important updates, or managing workflows across Google Workspace applications like Gmail, Docs, and Slides. It will also integrate with third-party services like Canva, OpenTable, and Instacart. Gemini Spark is currently being rolled out to a group of trusted testers today, with a wider beta version planned for Google AI Ultra subscribers in the United States next week.
Gemini 3.5 Pro: The Model That's Worth Waiting For
The May 19 launch also marks a notable absence. Google confirms it is actively working on Gemini 3.5 Pro, already used internally, with a deployment planned for next month. During the Google I/O 2026 keynote, Sundar Pichai told the audience: "I know you can’t wait to get your hands on it. Give us until next month to get it to you." (I know you can’t wait to get your hands on it. Give us until next month to get it to you.) according to information reported by Business Insider, which specifies that the news elicited audible reactions in the room.
On the most demanding pure reasoning benchmarks, Gemini 3.1 Pro remains for now the best option available in the Google range. The Pro version of the 3.5 model, expected in June 2026, should close this residual gap.
With Gemini 3.5 Flash, Google takes a symbolic step: fast models are no longer bargain alternatives but direct competitors to flagships. For general consumers, the change is seamless, with the Gemini app automatically switching to this new engine. For developers, the real interest lies in combining the model with Google Antigravity 2.0 and Managed Agents, which pave the way for agentic pipelines that were previously too costly or too slow to deploy at scale. The next step, with Gemini 3.5 Pro scheduled for June, will be to determine if Google can replicate this winning equation at an even higher level of reasoning.



No comments yet — start the discussion!