August 26, 2026
Our latest speech-to-text model, designed for accurate, intelligent real-time transcription.
Diego Melendo Casado
Senior Director of Engineering, Gemini Audio
Luke Leonhard
Chief of Staff, Gemini Audio, on behalf of the Gemini Audio team

Today, we’re introducing Gemini 3.5 Transcribe—our most accurate speech-to-text model yet, designed for intelligent voice interactions. Traditional speech recognition models often struggle with background noise, complex terminology, and the disfluencies of natural speech. Gemini 3.5 Transcribe, however, converts raw audio directly into accurate, polished, and properly formatted text.
Consumers are already benefiting from new voice features powered by this transcription model in products such as the Gemini app and Android, including Rambler on Android phones and the Gemini app for macOS. Now, developers can use the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform to build similar experiences with Gemini 3.5 Transcribe.
We built 3.5 Transcribe to fit seamlessly into developers’ workflows, whether they’re building voice agents, real-time captioning tools, or post-call analysis pipelines. The model is available through two separate APIs:
- Real-time streaming: Use
gemini-3.5-transcribe-livethrough the Live API for continuous, bidirectional streaming with latency of less than one second in interactive voice applications. - Pre-recorded audio processing: Use
gemini-3.5-transcribethrough the Interactions API to transcribe recordings, meetings, call logs, and more, with speaker attribution and word-level timestamps.
More accurate, intelligent transcription
Gemini 3.5 Transcribe is designed to capture the natural way people speak, helping it better understand user intent and recognize custom vocabulary so users can complete tasks with their voice.
- Intelligent transcription: Seamlessly handles self-corrections, such as “We’ll meet on Tuesday—no, Wednesday”; removes filler words such as “um” and “uh”; and automatically formats the text.
- Function calling: Through function calling, the model can delegate complex tasks, such as image generation and file analysis, to other Gemini models. This capability is currently available in the Gemini app for macOS.
- More accurate transcription: According to measurements by Artificial Analysis, it achieves an average Word Error Rate (WER) of 4.0% in streaming scenarios and 2.6% in non-streaming scenarios. The model also performs exceptionally well in noisy, real-world environments, accurately recognizing alphanumeric entities such as postal codes and order numbers.
- Custom vocabulary: Seamlessly adapts to user-provided custom vocabulary, recognizing technical terms and unusual spellings.
- Global language support: Automatically detects and transcribes more than 85 languages while flexibly handling regional accents and diverse dialects.
- Multi-speaker identification: Accurately distinguishes between speakers in pre-recorded audio and provides timestamped attribution for up to three speakers. Support for more than three speakers is still experimental.
Gemini 3.5 Transcribe delivers significant improvements over the previous transcription model, Chirp 3, with new capabilities, lower word error rates, and substantially lower latency. For example, according to measurements by Artificial Analysis, the time required to generate the final transcript has been reduced by 70%. On the FLEURS benchmark, which covers a range of widely spoken languages and regional settings, the model demonstrates accurate multilingual performance and improvements over Chirp 3, with WERs of 5.50% in streaming mode and 5.04% in non-streaming scenarios.


Experience intelligent transcription and advanced dictation
In addition to the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, 3.5 Transcribe goes beyond standard speech-to-text capabilities to make work across Google services feel more natural and intuitive. By bringing contextual awareness directly into everyday interfaces such as Gboard, Antigravity, the Gemini app, and Chrome, it can effortlessly capture linguistic nuances, user intent, and inline edits in text.
- In Gboard for Android, 3.5 Transcribe uses the new Rambler feature to turn spoken thoughts into well-formatted text while filtering out filler words. Users can also edit by voice, correct spelling mistakes, and change the writing style.
- In Google Antigravity, with the user’s permission, 3.5 Transcribe combines on-screen context and chat history to accurately transcribe content such as file names, agent reasoning, and the current document.
- In Google AI Studio, users can use 3.5 Transcribe in Build mode for instant vibe coding by voice.
- In the Gemini app for macOS, 3.5 Transcribe not only turns natural, conversational speech into clean, well-formatted text, but also seamlessly combines voice instructions with on-screen context to drive complex workflows. The model calls other Gemini models in the background to handle demanding tasks, allowing users to summarize local files, rewrite text across different applications, or generate images directly at the cursor—all with their voice.
- Chrome will support this capability soon: users will be able to dictate text by voice in any webpage input field, making it easy to dictate replies, compose posts, or prompt Gemini in Chrome by voice in a more natural and convenient way.
Early user feedback
With the Gemini Live API, developer platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents make it easy for developers to build and deploy high-performance voice-driven interfaces.