Skip to content
Main Site News Console

Introducing Gemini 3.5 Flash Cyber

· DeepMind Translated
DeepMind

July 21, 2026 Models

Raluca Ada Popa and Four Flynn

For years, Google has continuously invested in cybersecurity, pioneering automated vulnerability discovery to help protect codebases worldwide. Tools such as our code security agent CodeMender can automatically discover and fix critical software vulnerabilities. But as AI agents begin finding vulnerabilities faster than defenders can remediate them, addressing this global threat requires an approach that is powerful, affordable, and scalable.

Today, we are building on our longstanding efforts to help defenders prepare by introducing Gemini 3.5 Flash Cyber. Built on 3.5 Flash, this lightweight cybersecurity model has been fine-tuned to discover, validate, and patch vulnerabilities quickly and efficiently. It outperforms Gemini’s mainline Flash model on these tasks.

Flash’s performance and efficiency make it an ideal foundation for the cybersecurity models we are building. Built on Flash, 3.5 Flash Cyber provides a cost-effective, powerful alternative to larger and more expensive cybersecurity models.

Given the dual-use nature of this technology, we are taking a cautious approach to deploying 3.5 Flash Cyber. As part of a limited-access pilot program, 3.5 Flash Cyber will soon be made available exclusively through CodeMender to government agencies and trusted partners, with access expanding gradually over time. This will help frontline defenders get ahead of the threat by finding and fixing critical vulnerabilities before they are exploited, while reducing the risk of broader misuse.

We will also make CodeMender’s core capabilities, powered by general-purpose Gemini models, available directly to customers through the Gemini Enterprise Agent Platform.

The Search-Space Problem: Why Lightweight Models Excel at Code Security

Finding deep vulnerabilities requires exploring a vast execution search space. Relying on a single expensive large language model call can create a bottleneck. 3.5 Flash Cyber is particularly well suited to vulnerability discovery because these tasks require agents to scan large codebases and analyze a substantial number of code paths.

CodeMender calls 3.5 Flash Cyber multiple times, allowing the agent to analyze far more code paths and thereby discover and validate vulnerabilities. Multiple sub-agents then produce a high-quality consolidated report.

Thanks to its speed and low cost, 3.5 Flash Cyber can be easily integrated into high-frequency scans, time-sensitive release processes, or large-scale commit-scanning pipelines.

3.5 Flash Cyber Benchmark Results: An Efficient Alternative to Large Cybersecurity Models

We tested 3.5 Flash Cyber across a range of benchmarks. In particular, we evaluated it on the CyberGym benchmark, which assesses an AI agent’s ability to address hundreds of real-world software vulnerabilities. By taking advantage of 3.5 Flash Cyber’s low cost and configuring CodeMender to call 3.5 Flash Cyber up to five times for a single final report, the overall agent performed competitively with much larger models on CyberGym*.

CyberGym Success Rate (pass@1)

*Competitor results are scores self-reported by their providers

We also stress-tested the model’s capabilities outside CyberGym without safety mitigations. The Google Big Sleep team independently developed an evaluation focused on the model’s ability to find critical, subtle vulnerabilities in some of the world’s most complex codebases, including Chrome and Safari. In this evaluation, 3.5 Flash Cyber significantly outperformed both mainline 3.5 Flash and 3.6 Flash.

Big Sleep Evaluation Success Rate (pass@1)

We also evaluated 3.5 Flash Cyber on Google Chrome’s production commit-scanning pipeline. Because these vulnerabilities have not yet been publicly disclosed, this benchmark is not affected by contamination from Gemini or competitor model training data.

The results show a significant performance improvement for 3.5 Flash Cyber compared with 3.5 Flash. Note: Newer competitor models released after Opus 4.6 refused to perform these tasks because of their built-in safety mitigations and are therefore not shown here.

Chrome Production Commit-Scanning Pipeline Success Rate (pass@1)

In addition, 3.5 Flash Cyber consistently discovered more unique vulnerabilities than both mainline 3.5 Flash and Claude Opus 4.6. When tested on the highly complex V8 JavaScript engine with a fixed number of calls, 3.5 Flash Cyber found 55 confirmed unique issues, compared with 47 found by mainline 3.5 Flash and 36 found by Opus 4.6. Ten of these issues were not found by either of the other two models tested.

Basic cybersecurity models can get stuck in loops, repeatedly discovering the same issue while missing critical vulnerabilities. More capable models can expand the search space and find more unique issues.

As the number of calls increased, we found that 3.5 Flash Cyber continued to discover new code paths and vulnerabilities.

Real-World Applications at Google and Scaling Defense

Benchmarks are only part of the overall picture. 3.5 Flash Cyber in CodeMender has already discovered and fixed vulnerabilities in Google’s internal codebases, including projects across Chrome, Android, Cloud, Ads, and YouTube.

The discovery speed enabled by the lightweight model is already having a measurable impact.

For example, the Google Cloud Vulnerability Research team used 3.5 Flash Cyber to proactively protect our systems at record speed. In just two hours, the model discovered a remote code execution vulnerability in a public API and a memory corruption vulnerability in a sensitive production service. It then generated a 100%-reliable exploit for the remote code execution vulnerability that was able to bypass standard mitigations such as Address Space Layout Randomization (ASLR) and Write XOR Execute (W^X).

Early feedback from Wiz and Cloud CISO security engineering testers confirms that 3.5 Flash Cyber represents a significant capability improvement over the mainline 3.5 Flash model.

Empowering Defenders at Scale

Google’s leadership in software security gives us unique advantages. For example, the vulnerability database OSV.dev, operated by Google, contains more than 700,000 open-source vulnerabilities. In addition, more than 10 years of OSS-Fuzz results have helped us identify the highest-quality vulnerabilities.

This enables us to go beyond synthetic cybersecurity examples and teach the model how real security professionals work. Our model has learned to operate industry-standard tools, read millions of lines of code in large projects such as Chromium, and independently handle complex security tasks that require hours of sustained, in-depth analysis.

By using 3.5 Flash Cyber to power CodeMender, we are providing a powerful, scalable, and affordable architecture designed to help more defenders protect software.