August 28, 2026 · Google DeepMind
Google DeepMind Launches Gemini 3.5 Transcribe, Reporting 2.6% Word Error Rate Across 85+ Languages
My take: Google DeepMind launched Gemini 3.5 Transcribe this week, a speech-to-text model that reports a 2.6% average word error rate in standard mode and 4.0% in streaming, according to measurements by independent benchmarking firm Artificial Analysis. It auto-detects over 85 languages, removes filler words, separates speakers, and is 70% faster than its predecessor, Chirp 3.
For content creators, journalists, podcasters, and any professional who works with audio, this is a concrete and measurable tool: accurate transcription in multiple languages, without needing to manually correct the most common errors. I use transcription tools every day, and the difference between a 5% and a 2.6% error rate isn't minor — it's the difference between a quick review and a full edit.
The practical question is: how many hours a week do you spend transcribing, reviewing, or reformatting audio or video? If the answer is more than two, tools like this start to make economic sense. AI won't get it perfect every time, but at a 2.6% error rate, it's already useful most of the time.
Want to use these tools? See the unbiased reviews or back to the news.