Google introduces Gemini 3.5 Transcribe speech-to-text model
On August 26, 2026, Google DeepMind introduced Gemini 3.5 Transcribe, its latest speech-to-text model designed for precise and intelligent real-time transcription. The model converts raw audio into accurate, polished, formatted text, handling background noise, complex jargon, and disfluency cleanup. It offers multi-speaker attribution and word-level timestamps, and supports live language switches. Compared to its predecessor Chirp 3, it improves time to final transcription by 70% and achieves a 5.50% word error rate (WER) in streaming mode and 5.04% in non-streaming mode on the FLEURS benchmark. The model is available via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform, and is integrated into Google surfaces like Gboard, Antigravity, the Gemini app, and Chrome. Developer platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents support building voice-driven interfaces with it. Companies like vivo, Intellitek Health, and Lingopal have provided positive feedback.
What we know
It is available via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.
▤ 1 sources›
It achieves a 5.50% WER in streaming mode and 5.04% in non-streaming mode on FLEURS.
▤ 1 sources›
The model is a speech-to-text model with improved accuracy and latency compared to Chirp 3.
▤ 1 sources›
It is integrated into Google products like Gboard, Antigravity, the Gemini app, and Chrome.
▤ 1 sources›
Google DeepMind announced Gemini 3.5 Transcribe on August 26, 2026.
▤ 1 sources›
It supports multi-speaker attribution and word-level timestamps.
▤ 1 sources›
Time to final transcription improves by 70%.
▤ 1 sources›
Open any source to inspect its original language, when DoseFix received it, and the claims it supports.
Intelligent transcription with Gemini 3.5 Transcribe
deepmind.google · EN · Published · ReceivedLive reports
View allComments 0
Discuss this event in persistent threads. Live chat remains separate.
No comments yet. Start the conversation.