← Back

products

Google Launches Gemini 3.5 Transcribe for Real-Time Voice AI

Google has introduced Gemini 3.5 Transcribe, a speech-to-text model capable of supporting over 85 languages with real-time and offline processing modes, offering low latency and high accuracy.

AS1 News

voice-aitranscriptionmultilingualreal-time
REAL$0.0751+2.65%GOOGL$338.50-1.16%

Google has unveiled Gemini 3.5 Transcribe, a new AI model designed for automatic speech transcription. The model supports more than 85 languages and features both streaming and offline processing modes. In streaming mode, it achieves a latency of less than one second, making it suitable for real-time applications such as voice assistants and call centers. The model demonstrates error rates below 4% during streaming and 2.6% when processing recorded audio.

Gemini 3.5 Transcribe includes functionalities like automatic language detection, speaker diarization, and the removal of filler words, enhancing its utility in diverse environments. These features aim to improve the accuracy and clarity of transcriptions, facilitating better communication and data analysis in enterprise settings.

This release underscores Google's ongoing efforts to advance voice AI technology, providing developers and businesses with a robust tool for real-time speech recognition across multiple languages.

neutral

The new model enhances speech transcription capabilities, potentially improving voice-based applications and services across various industries.