Skip to content
Tools · Aug 26, 2026

Google DeepMind releases Gemini 3.5 Transcribe for real-time and pre-recorded audio transcription

A new speech-to-text model designed for precise, intelligent transcription with sub-second latency and support for over 85 languages.

Trust79
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • Gemini 3.5 Transcribe is a new speech-to-text model from Google DeepMind designed for real-time and pre-recorded audio transcription.

Google DeepMind introduced Gemini 3.5 Transcribe, a speech-to-text model optimized for precise and intelligent transcription in both real-time and pre-recorded audio scenarios. The model is designed to handle background noise, complex jargon, and disfluencies, converting raw audio into polished, formatted text.

Gemini 3.5 Transcribe is available via two APIs: a real-time streaming API with sub-second latency for interactive voice applications, and a pre-recorded audio processing API that supports speaker attribution and word-level timestamps. The real-time API uses the endpoint gemini-3.5-transcribe-live, while the pre-recorded API uses gemini-3.5-transcribe.

The model achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases, as measured by Artificial Analysis. It also demonstrates strong performance in noisy environments and accurately captures alphanumeric entities like postal codes and order IDs.

Gemini 3.5 Transcribe supports automatic detection and transcription for over 85 languages, including regional accents and dialects. It can attribute speech in pre-recorded audio to up to three speakers, with experimental support for more than three speakers. The model also handles live language switches and seamless streaming transcription.

Compared to its predecessor, Chirp 3, Gemini 3.5 Transcribe improves transcription latency by 70% and delivers better word error rates. On the FLEURS benchmark, it achieves a 5.50% WER in streaming mode and 5.04% WER in non-streaming use-cases across a set of top languages and locales.

Developers can access Gemini 3.5 Transcribe through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. The model is also integrated into Google products like the Gemini app, Android, macOS, Gboard, Antigravity, and Chrome to enable context-aware transcription and advanced dictation features.

Sources
  1. 01Google DeepMind — BlogIntelligent transcription with Gemini 3.5 Transcribe
Also on Tools

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.