July 21, 2026
Our latest Gemini models deliver efficient token usage, low latency, and reliable performance to help developers build AI agents at scale.

When building production-grade AI agents, developers and customers need greater token efficiency, lower latency, and more reliable performance. Our Flash series is designed to strike the optimal balance between efficiency and quality, enabling agent workflows to scale. Building on Gemini 3.5 Flash, we are introducing the following new Gemini models:
- 3.6 Flash: Our flagship model, offering improved performance in coding, knowledge work, and multimodal tasks. According to the Artificial Analysis Index, it uses 17% fewer output tokens than 3.5 Flash. On select benchmarks such as DeepSWE from Datacurve, we observed reductions of up to 65%, while also lowering the cost per output token.
- 3.5 Flash-Lite: The fastest and most cost-effective model in the 3.5 series. According to the Artificial Analysis Index, it generates 350 output tokens per second and significantly outperforms previous generations of Flash-Lite models in agent workflows.
- 3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of models and agent infrastructure. We are introducing a new, highly efficient specialized cybersecurity model, paired with the CodeMender code security agent, to deliver competitive, frontier-level performance.
In addition to the models announced today, Gemini 3.5 Pro is currently being tested with partners, and we plan to make it available to a broader audience as soon as it is ready. Meanwhile, our teams are already focused on developing the next generation of models. We have begun our most ambitious pretraining effort to date—Gemini 4—and are excited about the progress so far.
3.6 Flash: More efficient and higher quality than 3.5 Flash
Gemini 3.6 Flash was built directly in response to feedback from developers and customers using 3.5 Flash. In addition to improving coding and knowledge work capabilities, 3.6 Flash delivers substantially greater token efficiency. For example, on the Artificial Analysis Index, we observed that 3.6 Flash uses 17% fewer output tokens than 3.5 Flash. It also requires fewer reasoning steps and tool calls to complete multistep workflows.
This efficiency comes with a lower price than 3.5 Flash. 3.6 Flash is priced at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, reducing the total cost of each agent task and making agents more cost-effective to build and run.
3.6 Flash demonstrates greater token efficiency and produces fewer redundant outputs than 3.5 Flash on the OSWorld verification task (API).
Despite its greater efficiency, 3.6 Flash delivers improved performance over 3.5 Flash across a range of applications:
- On DeepSWE, 3.6 Flash achieves higher accuracy by reducing unnecessary code modifications and execution loops (49% vs. 37%). It also shows significant improvements in machine learning research on MLE Bench (63.9% vs. 49.7%).
- On OSWorld-Verified, its computer-use capabilities improve to 83.0%, compared with 78.4%. Computer use is now available as a built-in client tool through the Gemini API and Gemini Enterprise.
- In knowledge work, 3.6 Flash outperforms 3.5 Flash, achieving scores of 1421 vs. 1349 on benchmarks such as GDPval-AA v2. Customers including Hebbia and Harvey have found it particularly capable at multimodal tasks such as document analysis, chart and data analysis, and report drafting.


Customers report that 3.6 Flash improves both cost and quality, balancing token efficiency, accuracy, and speed in complex workflows and knowledge-intensive tasks:




Built with safety at the foundation
3.6 Flash launches with enhanced Frontier Safety protections covering chemical, biological, radiological, and nuclear (CBRN) domains, as well as cyberattack misuse scenarios. These protections significantly improve the model’s ability to resist jailbreak attacks. At the same time, the model has been trained to minimize refusals in beneficial use cases.
For more information, see the 3.6 Flash model card.
3.5 Flash-Lite: Built to scale agent workflows
In addition to the Flash series, we are releasing Gemini 3.5 Flash-Lite, designed for low-latency tasks and developer workflows where high throughput is critical, such as agent search and document processing.
3.5 Flash-Lite is the fastest model in the 3.5 series. According to measurements from Artificial Analysis, it generates 350 output tokens per second. 3.5 Flash-Lite is priced at $0.30 per 1 million input tokens and $2.50 per 1 million output tokens, while delivering substantially higher quality than 3.1 Flash-Lite. For developers and customers running high-throughput production workloads, 3.5 Flash-Lite offers excellent value.
3.5 Flash-Lite delivers lower latency than 3.5 Flash when executing high-volume tasks.
3.5 Flash-Lite enables agent systems to scale efficiently. Across different reasoning levels, the model significantly outperforms 3.1 Flash-Lite. Depending on the workload, developers can configure the model to use minimal or low reasoning for high-volume tasks at low latency and cost, or enable higher reasoning levels to handle multistep sub-agent workloads. The model also now includes computer use as a built-in tool, enabling it to reliably support these agent tasks across a range of scenarios.
3.5 Flash-Lite delivers significant improvements in coding and agent tasks, achieving 54% vs. 31% on Terminal-Bench 2.1; in long-context tasks, it scores 72.2% vs. 60.1% on GDM-MRCR v2; and in real-world task execution, it scores 1140 vs. 642 on GDPval-AA v2.

In fact, 3.5 Flash-Lite outperforms 3 Flash on many agent and coding evaluations, including SWE-Bench Pro (54.2% vs. 49.6%).