Brief IA

Cohere Unveils Transcribe, a Revolutionary Open-Source Voice Model

💻 Code & Dev·Tom Levy·

Cohere Unveils Transcribe, a Revolutionary Open-Source Voice Model

Cohere Unveils Transcribe, a Revolutionary Open-Source Voice Model
Key Takeaways
1Cohere has launched Transcribe, an open-source voice model for transcription, supporting 14 languages.
2Transcribe outperforms competing models with an average word error rate of 5.42 according to Hugging Face.
3Cohere plans to integrate Transcribe into its North platform and make it available for free via an API.
💡Why it mattersTranscribe could transform voice transcription by democratizing access to advanced and accurate tools.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Cohere Unveils Transcribe, a Revolutionary Open-Source Voice Model

Cohere, a company specializing in artificial intelligence, has recently introduced an innovative voice model named Transcribe. This automatic speech recognition model is open-source, meaning it is accessible to everyone for various applications such as note-taking or speech analysis. This initiative marks a significant advancement in the field of voice transcription.

The Transcribe model stands out for its lightweight design, with a total of 2 billion parameters. This makes it particularly suitable for use with consumer-grade GPUs, allowing individual users or small businesses to host it themselves. Transcribe is currently capable of processing 14 languages, covering a wide linguistic range including English, French, German, Italian, Spanish, Portuguese, Greek, Dutch, Polish, Chinese, Japanese, Korean, Vietnamese, and Arabic.

According to Cohere, Transcribe outperforms several competing models such as Zoom Scribe v1, IBM Granite 4.0 1B, ElevenLabs Scribe v2, and Qwen3-ASR-1.7B Speech. On the Open ASR leaderboard from Hugging Face, Transcribe achieves an average word error rate (WER) of 5.42, which is lower than all other models evaluated on this platform. This exceptional performance highlights the accuracy and efficiency of Transcribe in the field of voice recognition.

In terms of accuracy, Cohere claims that Transcribe has achieved an average win rate of 61% in human evaluations. These assessments considered the accuracy, consistency, and usefulness of the transcriptions provided by the model. However, it is important to note that Transcribe has shown slightly lower performance in transcribing languages such as Portuguese, German, and Spanish compared to its competitors.

Transcribe is also remarkable for its ability to process 525 minutes of audio in one minute, which is a notable performance for a model in its category. This processing speed could revolutionize how businesses and individuals manage large amounts of audio data.

Cohere plans to integrate Transcribe into its enterprise agent orchestration platform, called North. Additionally, the model is made available for free via an API, facilitating its adoption by a wide range of users. Transcribe will also be accessible on Model Vault, the inference platform managed by Cohere, further enhancing its accessibility and usage.

The popularity of voice recognition models is rapidly expanding, particularly due to the growing demand for note-taking and dictation applications like Granola and Wispr Flow. This trend underscores the increasing importance of voice transcription across various sectors.

Earlier this year, Cohere informed its investors of its growth ambitions, forecasting an annual recurring revenue of $240 million by 2025. Cohere's CEO, Aidan Gomez, also hinted that the company might consider going public in the near future, which could mark a new milestone in the startup's evolution.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.