Technologies
Back
Artificial Intelligence & Machine Learning

Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe

Dev.to
Advertisement468 × 90
Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe

Google has expanded its developer suite with the release of Gemini 3.8 Live and Gemini 3.5 Transcribe, designed to enhance real-time, voice-first applications. Gemini 3.8 Live offers advanced speech-to-speech capabilities, including an 'Extended Thinking' model for complex, multi-step reasoning. Key features include asynchronous function calling, visual context awareness, and support for over 97 languages. Complementing this, Gemini 3.5 Transcribe provides high-precision, low-latency speech-to-text conversion with a 4.0% Word Error Rate, featuring automatic code-switching and custom vocabulary biasing. Both models are accessible via the Gemini API and integrate with major platforms like LangChain, LiveKit, and Vercel. These tools aim to streamline the development of conversational agents, call center solutions, and real-time audio analytics, offering developers a robust infrastructure for building sophisticated, context-aware voice experiences that handle both complex reasoning and accurate transcription.

This is a summary. Read the full article at the original source:

Dev.to
Advertisement468 × 90
Share
Artificial Intelligence & Machine Learning

Related stories

Advertisement970 × 250