AI News

A New Era of Speech Translation: Introducing Gemini 3.5 Live Translate

Gemini 3.5 Live Translate is Google latest audio model, delivering near real-time speech-to-speech translation in over 70 languages.

Shawn H. avatar

Reviewed by Shawn H. Founder, AITrustList

Last verified MethodologyAI tools are ranked by traffic signals, not paid placement.

A New Era of Speech Translation: Introducing Gemini 3.5 Live Translate

Twenty years ago, translation at Google began as one of our pioneering machine learning experiments to turn the science of language into the magic of human connection. That experiment has come a long way, with over a trillion words translated for billions of users across our products every month.

Today, Google is taking its next leap forward with the release of Gemini 3.5 Live Translate—our latest audio model designed for near real-time, fluid speech-to-speech translation.

Let's dive into the core breakthroughs, use cases, and how you can experience this state-of-the-art model.


Watch: Introducing Gemini 3.5 Live Translate

To see the model in action, watch the official Google demonstration below:


Technical Breakthroughs: Moving Beyond Turn-by-Turn

Traditional speech translation systems operate in a "turn-by-turn" fashion. You speak, pause, wait for the system to process the transcription, translate the text, and generate speech synthesis before the other person can understand. This lag results in disjointed conversations and awkward silence.

Gemini 3.5 Live Translate fundamentally changes this dynamic:

1. Continuous Streaming Translation

Instead of waiting for the speaker to finish their thought, Gemini 3.5 Live Translate processes and translates speech continuously. It dynamically balances the trade-off between waiting for context to improve translation quality and translating immediately to stay in sync. The output audio streams smoothly, staying just a few seconds behind the speaker without awkward pauses.

2. Auto-Detection of 70+ Languages

The model automatically detects and translates between over 70 languages without requiring manual configuration. Users can speak naturally, and the model handles multilingual inputs on the fly.

3. Preserving Intonation, Pitch, and Pacing

Speech is more than just words; it is how we say them. Gemini 3.5 Live Translate generates natural-sounding translated speech that preserves the speaker's original vocal qualities—including their intonation, pacing, and pitch. This makes the translated audio sound like the speaker themselves, but speaking a different language.

4. Noise Robustness

Whether on a busy street or in a loud coffee shop, Gemini 3.5 Live Translate is built to be robust against background noise, ensuring reliable translation in unpredictable, real-world environments.


Product Integrations: Where to Experience It

Google is rolling out Gemini 3.5 Live Translate across its consumer and enterprise ecosystems:

1. Google Translate App (Android & iOS)

The model is rolling out globally in the Google Translate app. By connecting a pair of headphones, users can experience a conversational translation session that mirrors the speaker's tone in real-time across 70+ languages.

Private 'Listening Mode' on Android

For Android users, Google is launching a new "listening mode." When you don't have headphones handy or want a private translation, you can simply hold your phone to your ear like a normal call. The translated audio stream will play privately through your phone's earpiece.

Here is a demonstration of Listening Mode in action, translating a Spanish guided tour to English:

Download/Watch the Listening Mode demo video


2. Google Meet: Multi-Language Video Conferencing

Speech translation in Google Meet will soon be powered by Gemini 3.5 Live Translate, introducing major improvements:

  • 70+ Languages: Upgraded from the previous limit of just five languages.

  • 2,000+ Combinations: Supports cross-language translation in over 2,000 combinations (no longer restricted to translating to/from English).

  • Instant Interface: An updated UI offers immediate access to live speech translation.

Watch Google Meet participants communicate effortlessly across English, Mandarin, and Swedish in real-time:

The feature launches in private preview for select Google Workspace business customers this month, with a broader rollout planned later this year.


3. Developer API: Build with Gemini Live API

For developers, the model is available in public preview via the Gemini Live API and Google AI Studio. The preview model is named: gemini-3.5-live-translate-preview.

Developers can leverage the model to facilitate live interpretation for multilingual calls, meetings, virtual classrooms, broadcasts, and more.

Industry Reception

Mason Adams, Developer Evangelist at Agora, shared his team's findings:

"We tested the Gemini 3.5 Live Translate model at Agora and in our opinion it provided SOTA results, with low latency and high accuracy that set a new bar for real-time translation."

Other platforms like LiveKit Agents and Fishjam (utilizing the Media over QUIC/MoQ protocol) have also demonstrated seamless integrations, showcasing scenarios where participants speak their native language in a shared audio room and hear each other live.


Safety: Watermarked with SynthID

To ensure safety and authenticity in AI-generated voice content, all audio generated by Gemini 3.5 Live Translate is watermarked using SynthID.

Developed by Google DeepMind, SynthID embeds an imperceptible watermark directly into the audio output. It is completely inaudible to human ears but remains detectable to validation tools, helping to identify AI-generated content and prevent misinformation.


Conclusion

Gemini 3.5 Live Translate marks a significant shift in how we approach translation. By turning speech translation into a continuous, natural, and tone-preserving process, Google is closer to its ultimate goal: turning language translation from a utility into a tool for genuine human connection.

If you are a developer or AI enthusiast, you can start building with the model today in Google AI Studio.