Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe

Google has expanded its developer suite with the release of Gemini 3.8 Live and Gemini 3.5 Transcribe, designed to enhance real-time, voice-first applications. Gemini 3.8 Live offers advanced speech-to-speech capabilities, including an 'Extended Thinking' model for complex, multi-step reasoning. Key features include asynchronous function calling, visual context awareness, and support for over 97 languages. Complementing this, Gemini 3.5 Transcribe provides high-precision, low-latency speech-to-text conversion with a 4.0% Word Error Rate, featuring automatic code-switching and custom vocabulary biasing. Both models are accessible via the Gemini API and integrate with major platforms like LangChain, LiveKit, and Vercel. These tools aim to streamline the development of conversational agents, call center solutions, and real-time audio analytics, offering developers a robust infrastructure for building sophisticated, context-aware voice experiences that handle both complex reasoning and accurate transcription.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Researchers have introduced Dream-RSI, a novel framework designed to facilitate recursive self-improvement in artificial intelligence agents by levera…
The Download: AI’s trillion-dollar gamble and OpenAI’s biology data bid
MIT Technology Review’s latest newsletter examines the massive financial stakes behind the AI industry. Finance professor Jessica Wachter highlights t…
Mistral AI Partners with Mozilla to Enhance Private, Multilingual Browsing
Mistral AI has announced a strategic partnership with Mozilla to integrate its advanced artificial intelligence models into the Firefox ecosystem. Thi…



