
Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, pushing deeper into the race to make AI conversations feel less like chatbot exchanges and more like live dialogue.
The new models are described in a Google DeepMind model card published on September 15. Google says Gemini 3.8 Audio is based on Gemini 3 Pro and is built for real-time dialogue across audio, images, video and text. The models support a context window of up to 128,000 tokens and audio and text outputs of up to 64,000 tokens.
Gemini 3.8 Live is designed for scale and cost efficiency, while Gemini 3.8 Live Extended Thinking is aimed at more complex tasks. Google says the models are distributed through the Gemini API, Gemini App, Google AI Studio, Google Cloud Vertex AI and Search Live, while the Extended Thinking version will also reach Workspace tools including Gmail, Docs and Keep.
The timing is not random. OpenAI, Google, Meta and smaller startups are all trying to make voice and video AI feel faster, more natural and more useful. Text chatbots are already familiar. The next contest is whether AI can listen, see, reason and answer in real time without long delays or awkward handoffs between separate systems.
Google’s model card says the new Gemini 3.8 Audio models are intended for continuous streams of audio, video and text, and can deliver spoken responses in real time. That matters for live tutoring, customer support, accessibility, meetings, coding help and agentic workflows where a user may not want to keep typing prompts.
The safety section is also worth noting. Google says Gemini 3.8 Audio does not show meaningful new capabilities or material performance increases compared with Gemini 3.7 Flash for its Frontier Safety Framework assessment, and therefore is not likely to reach tracked or critical capability levels. The model card still acknowledges general foundation model limitations such as hallucinations, occasional slowness and jailbreak risks.
TechBooky recently covered Nuance Labs’ $50M bet on more natural AI avatars. Gemini 3.8 Live is part of that wider shift. AI is moving from text boxes into spoken, visual and always-on experiences. The winner may not be the model that only answers best on benchmarks, but the one that feels most natural when people actually talk to it.







