Google has officially unveiled Gemini 3.5 Transcribe, its next-generation audio foundation model designed for intelligent voice processing. The specialized system converts unstructured speech into polished, publication-ready text with automated formatting.
Most notably, the architecture strips conversational disfluencies like verbal pauses, accidental repetitions, and awkward interruptions directly from audio.
Automated Disfluency Removal Polishes Spoken Audio
Unlike traditional transcription engines, the model dynamically filters filler phrases including spoken “ums” and “ahs” during capture. It also seamlessly resolves mid-sentence corrections without creating duplicate words or broken phrasing.
The engine interprets contextual meaning to rewrite false starts naturally while maintaining the speaker’s core intent and message.
Developers can also configure custom vocabulary profiles to recognize technical jargon, company names, and abbreviations. The model supports over 85 languages with native accent adaptation across multiple dialects.
Lower Latency and Improved Recognition Benchmarks
Benchmark evaluations show Gemini 3.5 Transcribe achieves a 4.0% Word Error Rate during real-time streaming sessions. Non-streaming pre-recorded audio transcription records an even lower error rate of 2.6%.
Google confirmed the model delivers a 70% faster time-to-final-transcription when compared against its previous generation Chirp 3 engine.
The platform also includes multi-speaker diarization capabilities capable of tracking up to three distinct participants. Every transcribed utterance receives accurate word-level timestamps to simplify automated meeting indexing and review.
Built-in function calling allows the model to trigger downstream actions like image generation or summarization tasks.
Ecosystem Integration Powers Mobile and Desktop Apps
Google is deploying the transcription engine across consumer ecosystems, starting with the Gemini app for macOS. The model also powers the new Rambler voice dictation feature inside Gboard for Android.
Engineers are also preparing deep integration into the Chrome web browser to enable hands-free voice field entry.
By integrating smart disfluency filtering into transcription, Google transforms voice input into an efficient communication interface. This clean conversational pipeline significantly challenges rivals like OpenAI in enterprise audio agent deployment.