Overview
- Google and DeepMind announced SL2T will power sign-to-text in Gboard and Live Transcribe on the Pixel 11 family, initially translating American Sign Language into English.
- The system uses an on-device computer vision step to track 130 points across the face, hands, and torso, immediately discards the raw video, and sends only landmark coordinates to Google’s servers for translation.
- SL2T was trained on more than 100,000 hours of sign-language data across 50-plus languages with about one-quarter of the material in ASL, and it was taught to handle one-handed and left-handed signing for phone use.
- Google and its advisory committee warn the model has clear limits: it can miss regional signs, slang, complex ASL grammar, rapid fingerspelling, facially conveyed meanings, and can produce 'ghost text,' so it should not replace qualified interpreters.
- The company frames SL2T as an assistive input for low-stakes tasks like searches, messages, and short replies, says it does not keep logs unless users opt into studies, and plans wider device and language support at no extra cost.