Google DeepMind releases Gemini 3.5 Transcribe for real-time and pre-recorded audio transcription
A new speech-to-text model designed for precise, intelligent transcription with sub-second latency and support for over 85 languages.
1 source · cross-referenced
- Gemini 3.5 Transcribe is a new speech-to-text model from Google DeepMind designed for real-time and pre-recorded audio transcription.
Google DeepMind introduced Gemini 3.5 Transcribe, a speech-to-text model optimized for precise and intelligent transcription in both real-time and pre-recorded audio scenarios. The model is designed to handle background noise, complex jargon, and disfluencies, converting raw audio into polished, formatted text.
Gemini 3.5 Transcribe is available via two APIs: a real-time streaming API with sub-second latency for interactive voice applications, and a pre-recorded audio processing API that supports speaker attribution and word-level timestamps. The real-time API uses the endpoint gemini-3.5-transcribe-live, while the pre-recorded API uses gemini-3.5-transcribe.
The model achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases, as measured by Artificial Analysis. It also demonstrates strong performance in noisy environments and accurately captures alphanumeric entities like postal codes and order IDs.
Gemini 3.5 Transcribe supports automatic detection and transcription for over 85 languages, including regional accents and dialects. It can attribute speech in pre-recorded audio to up to three speakers, with experimental support for more than three speakers. The model also handles live language switches and seamless streaming transcription.
Compared to its predecessor, Chirp 3, Gemini 3.5 Transcribe improves transcription latency by 70% and delivers better word error rates. On the FLEURS benchmark, it achieves a 5.50% WER in streaming mode and 5.04% WER in non-streaming use-cases across a set of top languages and locales.
Developers can access Gemini 3.5 Transcribe through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. The model is also integrated into Google products like the Gemini app, Android, macOS, Gboard, Antigravity, and Chrome to enable context-aware transcription and advanced dictation features.
- Aug 26, 2026 · TechCrunch — AI
Robotics startup Generalist raises $200M extension at $3B valuation
Trust74 - Aug 26, 2026 · TechCrunch — AI
Voice AI startup Ringg raises $10M Series A extension led by Peak XV
Trust79 - Aug 25, 2026 · Hugging Face
Quantization-aware healing yields a 4-bit LLM that outperforms its full-precision original on most benchmarks
Trust79