Google has released Gemini 3.5 Transcribe, a speech-to-text model: The wider industry impact

Google has released Gemini 3.5 Transcribe, a speech-to-text model: The wider industry impact

Gemini 3.5 Transcribe removes filler words as you speak, and Chrome voice typing is next.

Google has released Gemini 3.5 Transcribe, a speech-to-text model that does not simply write down what you said. The model replaces Chirp 3, Google’s transcription engine from 2025, and the biggest gain is speed rather than accuracy. Time to final transcription improves by 70 percent, as measured by Artificial Analysis. Google cites a 4 percent word error rate on streaming audio and 2.6 percent on pre-recorded files, while on the multilingual FLEURS benchmark the figures are 5.50 percent streaming and 5.04 percent non-streaming, against 7.32 percent for Chirp 3 on live speech. Language coverage stretches past 85, with automatic detection and support for regional accents, which matters in a market where switching languages mid sentence is routine. Developers get two endpoints for this: gemini-3.5-transcribe-live through the Live API for sub-second streaming, and gemini-3.5-transcribe through the Interactions API for meetings and call logs.

Say “let’s meet Tuesday, no, Wednesday,” and it keeps only Wednesday. It tidies you up while you talk. Filler words disappear. It formats text as it goes, and it will adapt to a custom vocabulary list if your work involves jargon that trips up ordinary dictation. Accuracy moves less dramatically. For recorded audio, the model attributes speech to as many as three people and adds word-level timestamps. Anything beyond three speakers is still experimental.

Rambler, the Gboard feature currently restricted to the Pixel 11, runs on it, and Google says it will reach more Gemini Intelligence devices later this year. Plenty of people have already used this model without knowing it. Google Antigravity and AI Studio’s Build mode are live for developers, and Chrome support is promised soon, which will bring dictation to any web field. The tradeoff deserves a moment. A model that rewrites you is useful for a WhatsApp reply and risky for an interview quote or a legal note. Cleaner text is not the same thing as accurate text. Get the latest technology news and updates. Download the TOI App.

The Gemini app on macOS gets it from today, where function calling lets a spoken command hand off image generation or file analysis to other Gemini models.

One thought on “Google has released Gemini 3.5 Transcribe, a speech-to-text model: The wider industry impact

  1. Pingback: Google Cloud has unveiled Gemini Enterprise for Legal, an AI: The wider industry impact

Leave a Reply

Your email address will not be published. Required fields are marked *