This advanced press report was prepared to cover the behind-the-scenes details of Google launching AI sign language translation technology and related topics in the global technology and regulatory sector.
Table of contents:
- Integrating sign language translation into Pixel 11 phones
- How the model works and benchmark analysis
- Direct translation and surpassing traditional standards
- Governance, regulatory constraints, and principles of use
- Technical limitations and future expansion aspirations
- Frequently asked questions
Integrating sign language translation into Pixel 11 phones
Google DeepMind has officially announced the launch of its AI model dedicated to translating sign language into text, named SL-to-Text. This marks the first time an AI feature specifically for sign language has shipped inside a consumer smartphone application. The model supports sign-to-text dictation via the Google keyboard and the Live Transcribe app on the new Pixel 11 smartphone series, starting with American Sign Language (ASL) translated into written English. This unique feature allows deaf and hard-of-hearing users to point their phone camera and sign wherever they are used to typing—whether searching the web, drafting messages and emails, querying Gemini, or even responding instantly in face-to-face conversations via the Live Transcribe app. Google noted that testers found using sign language faster and more natural compared to traditional English typing.
How the model works and benchmark analysis
The SL-to-Text model was trained on more than 100,000 hours of diverse sign language data covering over 50 sign languages globally, with ASL making up about a quarter of that data. Rather than processing raw video—which consumes heavy processing power and poses privacy risks—the system relies on an on-device local model called MediaPipe Holistic to track 130 key landmarks on a person’s face, body, and hands, then sends only the resulting geometric coordinates to Google’s advanced servers while immediately and completely discarding the original camera feed to prevent any violation of user privacy and ensure strict security.
Direct translation and surpassing traditional standards
The model translates directly from body pose coordinates and landmark points into written text, bypassing the intermediate gloss annotations that previous systems relied on. DeepMind emphasizes that older glosses failed to capture the rich, non-linear dimensions of sign language, such as facial expressions, spatial arrangements, and body movement. Direct translation from landmark points significantly raises translation quality as the volume of available data increases without imposed limits. On the ASL benchmark evaluation scale, the model achieved an exceptional score of 70 points, significantly outperforming any previous results recorded in this field.
Governance, regulatory constraints, and principles of use
To ensure responsible use of the technology, Google established the Sign Language AI Advisory Committee, which includes prominent global organizations such as the National Association of the Deaf and the World Federation of the Deaf. The company’s impact report emphasized that the first version of the model is intended as an assistive tool for non-critical contexts only, explicitly ruling out the use of the feature in medical, legal, policing, or employment and health decision frameworks. The report also clarified that the feature is not considered a legal substitute for official obligations under the Americans with Disabilities Act in the interest of public safety.
Technical limitations and future expansion aspirations
Current technical limitations include reduced accuracy in low lighting, lack of support for multi-person simultaneous signing, and a capped 60-second time limit per clip with no cross-clip conversation memory. Furthermore, the model has not been trained on minors under 18 years of age. Notably, the project began as an initiative by a deaf engineer at Google named Sam Sepah, with deaf talent participating in every stage of data collection, development, and deployment. Google plans to expand the technology to new sign languages and additional devices at no extra cost to users, as active research into sign language generation continues.
Frequently asked questions
Question: What is the new technology launched by Google DeepMind for sign language?
Answer: Google launched the SL-to-Text model to translate sign language directly into written text via the Pixel 11 camera.
Question: How does the sign language translation model protect user privacy?
Answer: The model analyzes geometric landmark points of the face and hands and transmits only their coordinates while immediately deleting the original camera video.
Question: Can the feature be used in medical examinations or legal courts?
Answer: No, the feature is officially classified as an everyday assistive tool and is prohibited from use in medical, legal, policing, and official decision frameworks.
Question: What are the current temporal interaction limits and eligible user age?
Answer: The maximum clip length is 60 seconds, with a current focus on users over 18 years old to ensure model accuracy.