Skip to content
Main Site News Console

Introducing Gemini 3.6

· DeepMind Translated
DeepMind

July 21, 2026

Our latest Gemini models deliver the efficiency, low latency, and reliability needed to build large-scale AI agents.

When building production-grade AI agents, developers and customers require greater token efficiency, lower latency, and more reliable performance. Our Flash family of models is designed to strike the optimal balance between efficiency and quality, supporting the scaling of agentic workflows. Building on Gemini 3.5 Flash, we are introducing new Gemini models:

  • 3.6 Flash: Our workhorse model, offering improved performance in coding, knowledge work, and multimodal capabilities. According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash. On certain benchmarks like Datacurve’s DeepSWE, we observed reductions of up to 65%, with a lower cost per output token.
  • 3.5 Flash-Lite: Our fastest and most cost-effective 3.5-class model. According to the Artificial Analysis Index, it outputs 350 tokens per second while significantly outperforming the previous generation of Flash-Lite in agentic workflows.
  • 3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration between models and agentic infrastructure. We are introducing a combined solution: a brand-new, highly efficient, cybersecurity-focused model paired with our CodeMender code security agent, delivering competitive performance on frontier tasks.

In addition to today’s releases, Gemini 3.5 Pro is currently being tested with partners, and we plan to make it widely available as soon as it is ready. Meanwhile, our team is already focused on building the next generation of models. We have kicked off our most ambitious pre-training run to date—Gemini 4—and are excited about the progress.

3.6 Flash: More Efficient and Higher Quality than 3.5 Flash

Gemini 3.6 Flash was built directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only delivers improvements in coding and knowledge work but also significantly improves token efficiency. For example, on the Artificial Analysis Index, we see that 3.6 Flash consumes 17% fewer output tokens than 3.5 Flash. It also requires fewer reasoning steps and tool calls to complete multi-step workflows.

This efficiency boost comes alongside lower pricing compared to 3.5 Flash. Priced at $1.50 per million input tokens and $7.50 per million output tokens, 3.6 Flash reduces the overall cost of each agentic task, making it more cost-effective to build and run agents.

3.6 Flash is more token-efficient in OSWorld-Verified tasks and more streamlined than 3.5 Flash (API)

Even while being more efficient, 3.6 Flash delivers performance gains over 3.5 Flash across multiple use cases:

  • 3.6 Flash achieves higher accuracy with fewer unnecessary code modifications and fewer execution loops, as demonstrated by DeepSWE (49% vs. 37%); it also shows significant improvements in machine learning research, as shown by MLE-Bench (63.9% vs. 49.7%).
  • Its computer-use capabilities have improved, as shown by OSWorld-Verified (83.0% vs. 78.4%). Computer use is now available as a built-in client-side tool via the Gemini API and Gemini Enterprise.
  • It outperforms 3.5 Flash in knowledge work, as shown by benchmarks like GDPval-AA v2 (1421 vs. 1349). Customers like Hebbia and Harvey have found it particularly outstanding for multimodal tasks such as document parsing, chart and data analysis, and report writing.

Customers report that 3.6 Flash represents a major step forward in both cost and quality, balancing token efficiency, accuracy, and speed in complex workflows and knowledge-based tasks:

Built-in Safety

3.6 Flash is released with enhanced Frontier Safety safeguards, covering areas such as chemical, biological, radiological, and nuclear (CBRN) risks, as well as cyberattack abuse. These safeguards significantly increase the model’s resistance to jailbreak attacks. At the same time, the model has been trained to minimize false refusals for benign use cases.

For more information, please see the 3.6 Flash model card.

3.5 Flash-Lite: Built for Scaling Agentic Workflows

Alongside Flash, we are also releasing Gemini 3.5 Flash-Lite, designed specifically for low-latency tasks and developer workflows where high throughput is critical, such as agentic search and document processing.

3.5 Flash-Lite is the fastest model in the 3.5 family. According to measurements by Artificial Analysis, it runs at 350 output tokens per second. Priced at $0.30 per million input tokens and $2.50 per million output tokens, and offering significantly better quality than 3.1 Flash-Lite, it provides compelling price-to-performance for developers and customers running high-throughput production traffic.

3.5 Flash-Lite executes high-volume tasks with lower latency than 3.5 Flash.

3.5 Flash-Lite enables agentic systems to scale efficiently. It significantly outperforms 3.1 Flash-Lite across all thinking levels.