Google Gemini 3.5 Transcribe Tops GPT-Transcribe on 2.6% WER as OpenAI Undercuts on Price
Updated
Updated · KDnuggets · Sep 28
Google Gemini 3.5 Transcribe Tops GPT-Transcribe on 2.6% WER as OpenAI Undercuts on Price
3 articles · Updated · KDnuggets · Sep 28
Summary
Google’s Gemini 3.5 Transcribe posted a 2.6% word error rate for pre-recorded audio and 4.0% for streaming, outperforming OpenAI’s GPT-Transcribe benchmarked at 19.27% on Common Voice across 22 languages.
The comparison is unusually direct because both labs released new flagship transcription models within four weeks and split them into streaming and file-audio versions.
Gemini’s edge is built-in utility: speaker diarization for up to three speakers, word-level timestamps, custom vocabulary support and coverage for more than 85 languages, with Google also touting a 70% speed gain over Chirp 3.
OpenAI’s model is positioned more on cost and simpler deployment, priced at $0.0045 per minute for file transcription and $0.017 for streaming session audio, while still requiring separate models for diarization and timestamps.
The side-by-side suggests Gemini is better suited to meetings and call logs with multiple speakers, while GPT-Transcribe fits single-speaker transcription and live captioning where low cost and low-latency streaming matter most.
Why did OpenAI deliberately strip speaker tracking from its newest model, handing Google a massive advantage in the AI transcription war?
If pristine benchmark scores overstate AI accuracy, how will these 2026 transcription models actually handle real-world overlapping speech and heavy accents?