Google Integrates AI for Sign Language on Pixel
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
A Major Technological Advancement for Sign Language
Artificial intelligence has made significant strides in the realm of spoken languages over the past few decades. Tools like machine translation and voice recognition have become essential for hearing users. However, this technological revolution has not yet fully integrated sign languages, which encompass over 200 variants used by approximately 70 million deaf and hard-of-hearing individuals worldwide.
Today, a sign language-to-text translation model named SL2T marks a significant advancement. This massively multilingual model is now integrated into consumer products. It enables the dictation of sign language into text on devices such as Gboard and Live Transcribe, available on the Pixel 11. The model initially supports the translation of American Sign Language (ASL) into English, with plans to expand this functionality to other languages and devices.
This technology offers deaf users the ability to sign to conduct web searches, write messages or documents, and interact with virtual assistants like Gemini. In the Live Transcribe app, users can sign their responses during conversations, eliminating the need to type. According to early feedback, signing in ASL is perceived as faster and more natural than typing in English.
The Importance of Sign Languages
Sign languages play a central role in deaf communities around the world, constituting a key element of their cultural identity. There is great diversity among deaf individuals in terms of sign language proficiency, speech, reading, and writing skills. Therefore, it is crucial to support access to information in all these modalities. Sign language processing technology offers opportunities to enhance communication between deaf and hearing communities.
However, progress in this area has been slow, partly due to the complex challenges associated with creating AI for sign languages. Unlike speech transcription, which is a sequential mapping between sound and text, sign languages require true automatic translation. They have their own distinct grammars and lexicons, and understanding them involves following complex physical movements of the hands, arms, torso, head, and face.
Previous attempts, such as sign language gloves, have failed to capture the complexity of these languages. SL2T aims to overcome these limitations by providing advanced visual perception and comprehensive linguistic translation.
How SL2T Works
SL2T was developed by combining a user-centered approach with massive data scaling. The model was trained on over 100,000 hours of data covering more than 50 sign languages, with about a quarter in ASL. This diversity allows the model to learn common linguistic structures, thereby surpassing monolingual models.
To ensure user privacy, SL2T treats sign language as a sequence of pose landmarks rather than raw video footage. An on-device model, MediaPipe Holistic, tracks the position of points on the signer, and only these geometric coordinates are sent to the server for translation, eliminating the original video.
SL2T directly translates these coordinates into text, avoiding intermediate annotations known as "glosses." These glosses do not capture the rich, non-linear aspects of sign languages, such as non-manual markers and spatial constructions. Direct translation enhances quality by relying directly on the data.
According to key benchmarks like FLEURS-ASL, SL2T is the most performant sign language translation model to date. It achieves a score of 70 BLEURT in zero-shot, surpassing previous scores. However, optimizing academic benchmarks does not guarantee practical usability. That is why efforts have been made to minimize streaming latency, avoid errors on unsigned inputs, ensure fairness for left-handed signers, and improve performance for one-handed signing.
Collaboration with the Deaf Community
The development of SL2T has been guided by close collaboration with the deaf community. Deaf perspectives have influenced every stage of the project, from conceptualization by Sam Sepah, a deaf employee at Google, to data collection with deaf partners, and evaluation of the technology with deaf experts.
To ensure responsible deployment, Google has established the AI Sign Language Advisory Committee (AISLAC), which brings together global deaf organizations and experts. This participatory governance model allows the most affected communities to directly influence development priorities. A joint impact report has been drafted for the release of SL2T 1.0, detailing the current capabilities and limitations of the technology.
Looking Ahead
SL2T builds on decades of foundational research, but the introduction of ASL on phones is just the beginning. Google aims to make information universally accessible, which involves achieving full parity with spoken and written languages. The team is working to extend this technology to other sign languages, to sign language generation, and to advanced AI capabilities. Progress will be shared responsibly to make access via sign languages standard in the digital landscape.
Users can experience SL2T in Gboard and Live Transcribe on the Pixel 11, with other devices to follow, at no additional cost.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.