Google senior vice president James Manyika has published a feature article announcing Google’s latest milestone in advancing AI accessibility and multilingual support: Google’s technologies and products now support more than 300 languages worldwide, reaching as many as 7 billion people (about 86% of the global population). Manyika laid out the vision in Google’s essay on building AI for every language.
To further enable AI to comprehend the diverse ways people communicate across the real world, Google announced a departure from the traditional “speech-to-text-then-translate” framework, turning instead to native audio intelligence that processes speech directly (that is, letting AI understand what a user says firsthand). Simultaneously, it unveiled Gemini 3.5 Live Translate, Gemini 3.5 Transcribe, and TranslateGemma, an open-source model built for on-device computation. It even introduced sign-language recognition for the first time in the newly released Pixel 11 smartphone.
Farewell to Stilted Machine Voices: Gemini 3.5 Brings Native-Audio Two-Way Translation
Past speech-recognition systems mostly relied on a cumbersome multistep process: first transcribing audio into text, then processing the text and synthesizing speech. This approach, however, often stripped away the most authentic tone, rhythm, and emotion of human communication.
To address this, Google trained Gemini models to process native audio directly, launching two core speech technologies:
- Gemini 3.5 Live Translate: This new technology supports real-time speech translation across 70 languages and over 2,000 language-pair combinations. Its greatest breakthrough lies in naturally capturing a speaker’s emotional shifts and tone, and even seamlessly handling the multilingual, code-mixed conversations common in real life (such as Chinese-English mixing, Spanglish, or Hinglish).
- Gemini 3.5 Transcribe: As Google’s most accurate speech-to-text model to date, it can precisely convert raw audio into neatly formatted text even in noisy environments or amid dense technical jargon. This technology also brings a new feature called “Rambler” to the Android Gboard keyboard, which automatically removes filler words while speaking, corrects grammar and punctuation, and lets users edit and rewrite through voice commands.
Advancing Open Source and Local Computing: TranslateGemma Makes Offline Translation Possible
Addressing the plight of more than 3 billion people worldwide who face unstable connectivity or no internet access, Google this time also unveiled a series of lightweight open-source translation models built on the Gemini architecture: TranslateGemma.
TranslateGemma supports 55 languages in total. Its greatest advantage lies in running locally on-device with extremely high efficiency, meaning that even with no network connection whatsoever, users can still enjoy a high-quality translation experience.
Moreover, to serve feature-phone users in resource-scarce regions, Google has partnered with Viamo to test the “Ask Viamo Anything” (AVA) service, which supports a voice AI assistant, in Rwanda. It has already successfully handled more than 2 million voice queries.
Prioritizing Localization and Accessibility: Pixel 11 Debuts Sign-Language-to-Text
Beyond spoken translation, Google has taken an important step in accessible design. The company announced Sign Language-to-Text (SL2T) technology supporting more than 50 sign languages. This innovative feature will debut on the newest Pixel 11 phone, deeply integrated with the Gboard input method and Live Transcribe, initially supporting conversion between American Sign Language (ASL) and English.
In building its language databases, Google has turned toward local grassroots collaboration rather than relying solely on web scraping. Examples include the WAXAL project with African institutions (covering 27 sub-Saharan African languages), Project Vaani in India (which has amassed 109 local languages), and Language Explorer, an open-source interactive tool that visualizes data on more than 7,000 languages. All demonstrate Google’s resolve to realize its 1000 Languages Initiative.
From Literal Translation to Semantic Resonance, Edge Computing Completes the Puzzle
For a long time, the technology industry’s development has been dominated by a handful of powerful languages. The many updates Google has released this time are not merely a display of technological firepower but an important recalibration of the generative AI development path.
Traditional machine translation, which relies on text as an intermediary, is like viewing the world through frosted glass: one can make out the outline yet loses the true colors. Gemini 3.5 Live Translate’s ability to “hear” audio directly preserves the subtext and emotional resonance of human communication, which will greatly transform the future of cross-border meetings and cross-cultural exchange.
On another front, the launch of TranslateGemma strikes at another pain point of current AI development: cloud computing cost and network dependence. Bringing powerful translation capability down to run offline on-device not only enhances privacy protection but also truly pushes AI toward regions with lagging infrastructure.
When most tech giants pursue ever-larger model parameters, Google chooses to invest resources in preserving the local accents of minority languages, and even in visualizing sign language into text. This undoubtedly sets a deeply humane industry benchmark for “AI technology equity.”
Support Our Threat Intelligence
Find our tech and OS security coverage helpful? Support our work today and unlock a 100% ad-free reading experience!