The quick read
- A conversational agent must handle timing as well as words.
- Evaluate the languages, names, background noise and interaction patterns of the intended audience.
What is new
Google introduced Gemini 3.8 Live and its Extended Thinking variant for developers on September 15. The announcement describes speech-to-speech interaction with background tool calls and visual context. It also highlights Gemini 3.5 Transcribe, which had been released the previous month; the two models address different parts of an audio workflow.
Why architecture matters
A conversational agent must handle timing as well as words. It needs to respond appropriately to interruptions, pauses and corrections while making any external action clear. A transcription service instead needs an accurate, usable representation of what was said.
What to test
Evaluate the languages, names, background noise and interaction patterns of the intended audience. Check tool permissions and make consequential details reviewable. Treat vendor benchmark claims as a starting point, and measure the complete experience rather than only isolated transcription or response speed.
Sources & notes
AI-assisted editorial content checked against the linked sources.
blog.google — official reference
Sources reviewed for the September 2026 launch edition.
