Reasoning inside a voice loop: what Gemini 3.8 Live Extended Thinking leaves unanswered
On September 15 Google released two live models for the Gemini API: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The first is described as scalable and cost-efficient, with visual grounding, meaning it can work with what a camera sees while it talks. The second adds multi-step reasoning for harder tasks. Both are in the Gemini API and AI Studio, enterprise access is in private preview, and the same models power Search Live, Gemini Live and Workspace.
That is what the announcement gives us. Latency numbers, thinking prices and interruption behaviour are not in it, and those are the three things a voice agent lives or dies on. So this piece is about the decision you will have to make, and about which numbers to collect before you make it.
Thinking has a sound, and the sound is silence
In text, extended reasoning costs you seconds you can hide behind a spinner. In a voice loop there is no spinner. A person who asks a question and hears nothing for three seconds assumes the line dropped. Human turn-taking gaps are a few hundred milliseconds, and that expectation does not care how clever the model is.
So reasoning in a live loop is a scheduling problem more than a quality problem. You can talk while thinking (a filler phrase, an acknowledgement), you can answer with the fast model first and correct yourself later, or you can accept a pause and warn the caller. Each has a price in naturalness, and none of them is free.
A voice agent does not pay for thinking in tokens alone, it pays in dead air, and dead air is the more expensive currency.
Three budgets that fight each other
Latency, cost and interruptibility pull in different directions, and I would track all three from day one.
Latency is time from the user finishing a sentence to the first audible syllable, measured at the P50 and P95, because callers remember the slow ones. Cost is the bill per completed conversation, and here the source is silent on how Extended Thinking is charged, so I will not guess. What I can say is that reasoning tokens usually cost as output, and in a phone call the same task may run five, ten, twenty turns. Multiply accordingly.
Barge-in is the third, and the least discussed. When the caller interrupts mid-answer, the system has to stop audio, drop the unspoken remainder, and decide whether the interrupted thought still matters. With a reasoning model that is halfway through a chain of steps, an interruption can invalidate work you already paid for. Whether the live models cancel that work cleanly, and whether you are billed for it, is exactly the kind of thing the changelog does not tell you.
A routing rule instead of a single model
The sensible design, as I see it, is not to pick one of the two models. It is to route per turn. Most turns in a support or booking call are trivial: confirm a date, read back a number. They belong on the fast, cheap variant. A small share needs real reasoning, such as comparing two policies or resolving contradictory information, and only those should reach the Extended Thinking model, ideally with a spoken hold line while it works.
turn -> classify(complexity, user_waiting_tolerance)
simple -> live model, answer immediately
complex -> say a short holding phrase, call extended thinking
interrupt -> cancel pending reasoning, resume on the fast path
The classifier is the part that will cost you the most engineering. A wrong call in one direction gives a slow, expensive agent, in the other a confident wrong answer. I would start with crude rules (question length, presence of comparison words, tool calls needed) and log every decision so the rules can be tuned against real calls.
What to measure before a rollout
I haven't run these models myself, so treat everything above as a design argument, not a result. Before I let one near customers I would record, for each turn: time to first audio, whether the turn used thinking, tokens billed, and whether the user interrupted. Two weeks of that from a pilot gives you a cost per conversation and a real P95, which is more than any demo will.
The visual grounding is the part I am most curious about, honestly. An agent that sees the router's blinking lights while it talks the caller through a reset is a different product from a voice bot. Whether Google's pricing and latency let that run at scale is still an open question, and the public documentation I have seen does not settle it.