Gemini 3.5 Transcribe: Two APIs, Latency <1s, 85+ Languages

Google introduces Gemini 3.5 Transcribe, a voice recognition model designed for voice interactions. The offering is now available in public preview for developers and businesses, with public availability starting on macOS and Android, and set to expand to Chrome. The model claims improvements in accuracy and latency compared to Chirp 3 and supports over 85 languages.
Developers and Businesses Access Public Preview
Gemini 3.5 Transcribe is offered in public preview via the Gemini API in Google AI Studio and through Google Antigravity for developers. Businesses can access it in public preview via the Gemini Enterprise Agent platform, with availability soon to be announced in Gemini Enterprise for customer experience. For the general public, access is open in the Gemini app on macOS in English and via Rambler on Android in certain countries and languages. A launch on Chrome is expected soon.
Measured Accuracy and Improved Latency Compared to Chirp 3
According to Analyse Artificielle, the model achieves an average error rate of 4.0% in streaming and 2.6% in non-streaming scenarios, showing good performance in noisy environments, including with alphanumeric entities. Google presents these results as a significant advancement over Chirp 3, with new capabilities, better error rates, and substantially improved latency. Measured by Analyse Artificielle, the final transcription time improves by 70%. On the FLEURS benchmark, the model outperforms Chirp 3, showing an error rate of 5.50% in streaming mode and 5.04% in non-streaming mode.
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Two APIs Distinguish Streaming at <1 s and Pre-recorded Audio
The offering relies on two interfaces. The Live streaming API, consumed via the model gemini-3.5-transcribe-live, provides a continuous bidirectional stream for interactive voice applications, with latency under one second. For pre-recorded audio, the Interactions API, via the model gemini-3.5-transcribe, supports the transcription of recordings, meetings, and call logs, with speaker attribution and word-level timestamps.
The Model Adapts Transcription to Natural Speech Style
The model handles self-corrections, eliminates filler words, and automatically formats text. It captures the natural speech style to better understand intent and adapts to custom vocabulary while managing live language changes. Coverage extends to over 85 languages and takes into account accents and dialects. In pre-recorded audio, diarization is offered for up to three speakers, with support beyond that being experimental. The model can delegate tasks to other Gemini models via function calls, a capability available in the macOS app.
Integrations: Gboard, Antigravity, AI Studio, and macOS App
Google reports having already observed user benefits through the Gemini app and on Android with Rambler. On Gboard for Android, Rambler transforms speech into well-formatted text, filters out filler words, and allows editing, correcting, and changing voice styles. On Google Antigravity, the model leverages, with permission, screen context and conversation history to optimize transcription, including on file names, agent thoughts, and open documents. In AI Studio, 3.5 Transcribe is accessible in Build mode for coding real-time voice applications. In the Gemini app on macOS, beyond the transcription itself, voice commands align with screen context for complex workflows, calling other Gemini models to synthesize local files, reuse text across applications, or generate images directly at the cursor, all through voice alone.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.