⚡
Brief IA
›

Gemini 3.5 Transcribe: Two APIs, Latency <1s, 85+ Languages

🔬 Research·Tom Levy·

Gemini 3.5 Transcribe: Two APIs, Latency <1s, 85+ Languages

Gemini 3.5 Transcribe: Two APIs, Latency <1s, 85+ Languages
⚡
Key Takeaways
1Two APIs: bidirectional streaming with latency <1 s and pre-recorded processing with speaker attribution and timestamping
2Announced accuracy: 4.0% in streaming, 2.6% offline; +70% on final time according to Artificial Analysis
3Availability: public preview via AI Studio, Antigravity, and Enterprise Agent; macOS app (English) and Rambler Android, Chrome coming soon
💡Why it matters — The model becomes accessible in public preview with announced accuracy and latency metrics and integrates with several Google products, expanding voice usage for developers, businesses, and the general public.

Google introduces Gemini 3.5 Transcribe, a voice recognition model designed for voice interactions. The offering is now available in public preview for developers and businesses, with public availability starting on macOS and Android, and set to expand to Chrome. The model claims improvements in accuracy and latency compared to Chirp 3 and supports over 85 languages.

Developers and Businesses Access Public Preview

Gemini 3.5 Transcribe is offered in public preview via the Gemini API in Google AI Studio and through Google Antigravity for developers. Businesses can access it in public preview via the Gemini Enterprise Agent platform, with availability soon to be announced in Gemini Enterprise for customer experience. For the general public, access is open in the Gemini app on macOS in English and via Rambler on Android in certain countries and languages. A launch on Chrome is expected soon.

Measured Accuracy and Improved Latency Compared to Chirp 3

According to Analyse Artificielle, the model achieves an average error rate of 4.0% in streaming and 2.6% in non-streaming scenarios, showing good performance in noisy environments, including with alphanumeric entities. Google presents these results as a significant advancement over Chirp 3, with new capabilities, better error rates, and substantially improved latency. Measured by Analyse Artificielle, the final transcription time improves by 70%. On the FLEURS benchmark, the model outperforms Chirp 3, showing an error rate of 5.50% in streaming mode and 5.04% in non-streaming mode.

⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

Two APIs Distinguish Streaming at <1 s and Pre-recorded Audio

The offering relies on two interfaces. The Live streaming API, consumed via the model gemini-3.5-transcribe-live, provides a continuous bidirectional stream for interactive voice applications, with latency under one second. For pre-recorded audio, the Interactions API, via the model gemini-3.5-transcribe, supports the transcription of recordings, meetings, and call logs, with speaker attribution and word-level timestamps.

The Model Adapts Transcription to Natural Speech Style

The model handles self-corrections, eliminates filler words, and automatically formats text. It captures the natural speech style to better understand intent and adapts to custom vocabulary while managing live language changes. Coverage extends to over 85 languages and takes into account accents and dialects. In pre-recorded audio, diarization is offered for up to three speakers, with support beyond that being experimental. The model can delegate tasks to other Gemini models via function calls, a capability available in the macOS app.

Integrations: Gboard, Antigravity, AI Studio, and macOS App

Google reports having already observed user benefits through the Gemini app and on Android with Rambler. On Gboard for Android, Rambler transforms speech into well-formatted text, filters out filler words, and allows editing, correcting, and changing voice styles. On Google Antigravity, the model leverages, with permission, screen context and conversation history to optimize transcription, including on file names, agent thoughts, and open documents. In AI Studio, 3.5 Transcribe is accessible in Build mode for coding real-time voice applications. In the Gemini app on macOS, beyond the transcription itself, voice commands align with screen context for complex workflows, calling other Gemini models to synthesize local files, reuse text across applications, or generate images directly at the cursor, all through voice alone.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.