INDEPENDENT MINDS. INTELLIGENT COVERAGE.

AI, ONLY. ALL ANGLES.

Google updates its real-time audio building blocks

Google’s developer update separates live conversation from transcription while expanding its audio toolkit.

PromptWireGlobal2 min read2026-09-15
Microphone and speaker exchanging crisp speech waves through small developer console
Conceptual illustration for PromptWire.

In this story

The quick read

  • A conversational agent must handle timing as well as words.
  • Evaluate the languages, names, background noise and interaction patterns of the intended audience.

What is new

Google introduced Gemini 3.8 Live and its Extended Thinking variant for developers on September 15. The announcement describes speech-to-speech interaction with background tool calls and visual context. It also highlights Gemini 3.5 Transcribe, which had been released the previous month; the two models address different parts of an audio workflow.

Why architecture matters

A conversational agent must handle timing as well as words. It needs to respond appropriately to interruptions, pauses and corrections while making any external action clear. A transcription service instead needs an accurate, usable representation of what was said.

What to test

Evaluate the languages, names, background noise and interaction patterns of the intended audience. Check tool permissions and make consequential details reviewable. Treat vendor benchmark claims as a starting point, and measure the complete experience rather than only isolated transcription or response speed.

Sources & notes

AI-assisted editorial content checked against the linked sources.

blog.google — official reference

Sources reviewed for the September 2026 launch edition.

KEEP EXPLORING.