Gemini 3.8 Live makes voice agents worth a production pilot
Google has shipped voice models that keep conversing while tools and multi-step reasoning run in the background. The technology now deserves a bounded production pilot, but long-session economics still need disciplined measurement.
What shipped
Google made Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking generally available through the Gemini API on September 15. The base model targets low-latency dialogue; Extended Thinking can reason and call tools in the background while continuing the conversation.
Strong, not solved
Google reports an 82.6 score and first place for Extended Thinking on the Artificial Analysis Speech to Speech Index, plus 68.6% on τ-Voice. That warrants testing, but its 35.1% result on Sierra’s banking benchmark says complex workflows are still far from solved.
Session economics
Published audio rates are $0.005 per input minute and $0.018 per output minute. Yet the Live API processes retained context again on later turns, while proactive listening is permanently enabled for the 3.8 models, making long-session economics different from a short demo.
What we would do
Virtual Arc would start with one bounded workflow, a human escalation path and a measurable final state. We would use the base model for routine turns, reserve Extended Thinking for complex branches, and expand only after measuring reliability and cost per completed task.
Virtual Arc’s view is that Gemini 3.8 Live makes voice agents worth a production pilot, but not a platform rewrite. The meaningful change is not that the model sounds smoother; it is that the session can keep speaking while tools run and, with Extended Thinking, while multi-step reasoning continues in the background. That can remove the dead air that makes tool-using voice systems feel broken. But Google’s billing mechanics make the headline per-minute rates incomplete: retained session context is processed again on later turns, proactive listening is permanently enabled, and thinking tokens are billable output. We would therefore route routine turns to the base model, escalate only genuinely complex work to Extended Thinking, compress context aggressively, and judge the pilot on cost per correctly completed task, interruption recovery, wrong-tool calls and human handoff quality. We would also keep a thin provider adapter around the Live API. A leaderboard win is not sufficient evidence for accepting deep vendor lock-in in an operationally messy workflow.