Google recently announced the launch of Gemini 3.5 Transcribe, its newest speech-to-text model. The model is engineered to sharpen real-time transcription accuracy, reduce latency, and strengthen multilingual capability. Unlike conventional speech recognition systems, this model goes beyond simply converting speech into text. It also interprets the speaker’s underlying intent and automatically organizes and formats the resulting output.
Automatically Removing Filler Words and Correcting Speech
Gemini 3.5 Transcribe automatically handles pauses, self-corrections, and verbal filler. For instance, if a user says “See you tomorrow – no, wait, make that the day after,” the model outputs the corrected final version directly. Additionally, it simultaneously strips out meaningless filler sounds like “um” and “uh.” Users can also supply a custom vocabulary list. This allows the model to correctly recognize product names, technical terminology, and unusual spellings. As a result, it reduces transcription errors in contexts such as technical meetings, healthcare settings, and enterprise customer service.
The model supports automatic recognition and transcription across more than 85 languages. It features optimization tuned for regional accents and dialects. When processing pre-recorded audio, it can distinguish and attribute speech to up to 3 distinct speakers while providing word-level timestamps. However, support for more than 3 speakers currently remains in experimental status.
Streaming Transcription With Sub-1-Second Latency
Gemini 3.5 Transcribe is offered through both a real-time API and a pre-recorded API. The real-time mode delivers continuous, bidirectional streaming transcription through the Live API, with latency under 1 second. The pre-recorded mode is designed primarily for transcribing meeting recordings, phone call logs, and other previously recorded audio.
According to Google, the model achieves an average word error rate of just 4% in real-time transcription and 2.6% in non-real-time mode. Meanwhile, it generates final transcription results 70% faster than Google’s previous speech-to-text models.
Availability
The model is currently available to developers in public preview through the Gemini API platform. It is also accessible to users of Gemini for macOS. Within Gemini for macOS, users can directly speak commands to ask the AI to analyze local files, rewrite text, and generate images.
Google has also confirmed that this capability will extend to Chrome in the future, at which point users will be able to compose and edit text directly within web input fields using natural spoken language.
Support Our Threat Intelligence
Find our zero-day alerts and CVE reports helpful? Support our work today and unlock a 100% ad-free reading experience!