← Gautam Parab

Gemini's Voice-Agent Lead Is 0.7 Points, Over a Rival Scored on One Trial

On 15 September Google released two voice models under one name. Gemini 3.8 Live is the fast, cheap one. Gemini 3.8 Live Extended Thinking reasons while it talks. The announcement credits the second with first place on Artificial Analysis’ Speech to Speech Index at 82.6, 68.6% on τ-Voice and 35.1% on what it calls “Sierra’s τ-Voice-banking benchmark.”

I went through the evaluation document Google links from its model card, and then through the leaderboards that document cites. Three things came out of that reading. The headline numbers belong to one of the two models. Its lead over OpenAI is smaller than the noise in the measurement. And the benchmark family behind the agent scores was designed around a reliability metric that nobody in this release reports.

Two models, one name

Google’s methodology PDF charts both models side by side. On τ-Voice, the Extended Thinking model scores 68.6%. The base Gemini 3.8 Live scores 30.1%, lower than the 37.7% Google’s own previous-generation Gemini 3.1 Flash Live reached at its high setting. The same PDF says the τ-Voice runs used “thinking high” for both model IDs, so on Google’s account the gap isn’t down to one of them being run at a lower setting.

The distinction matters because of where each model ships. The base model is the one rolling out “for everyone” in Search Live, and it is priced for that job: Artificial Analysis lists it at $0.84 per hour of input audio, against $3.50 for Extended Thinking and $5.83 for OpenAI’s GPT-Live-1 Astra.

τ-Voice task completion against cost per hour of input audio A scatter plot of six speech-to-speech models, with cost per hour of input audio on the horizontal axis and τ-Voice task completion on the vertical axis, both from Artificial Analysis as charted in Google's September 2026 evaluation document. Gemini 3.8 Live: $0.84, 30.1%. Gemini 3.1 Flash Live (Minimal): $1.50, 26.2%. Gemini 3.1 Flash Live (High): $1.75, 37.7%. Gemini 3.8 Live Extended Thinking: $3.50, 68.6%. Grok Voice Think Fast 2.0 (High): $4.80, 56.5%. GPT-Live-1 Astra (Medium): $5.83, 67.9%. The two Gemini 3.8 models are drawn as filled dots; the cheap one completes less than half as many tasks as the expensive one. τ-VOICE TASK COMPLETION VS COST PER HOUR OF INPUT AUDIO · ARTIFICIAL ANALYSIS, SEPT 2026 0%20%40%60%80% $0$2$4$6 Cost per hour of input audio (40-question Big Bench Audio subset) Gemini 3.1 Flash Live (Minimal): $1.50 per hour, 26.2% on τ-Voice — Artificial Analysis via Google, 15 Sept 2026 Gemini 3.1 Flash Live (High): $1.75 per hour, 37.7% on τ-Voice — Artificial Analysis via Google, 15 Sept 2026 Grok Voice Think Fast 2.0 (High): $4.80 per hour, 56.5% on τ-Voice — Artificial Analysis via Google, 15 Sept 2026 GPT-Live-1 Astra (Medium): $5.83 per hour, 67.9% on τ-Voice — Artificial Analysis via Google, 15 Sept 2026 Gemini 3.8 Live: $0.84 per hour, 30.1% on τ-Voice — Artificial Analysis via Google, 15 Sept 2026 Gemini 3.8 Live Extended Thinking (High): $3.50 per hour, 68.6% on τ-Voice — Artificial Analysis via Google, 15 Sept 2026 Gemini 3.8 Live · $0.84 · 30.1% Gemini 3.1 Flash Live (Minimal) · $1.50 · 26.2% Gemini 3.1 Flash Live (High) · $1.75 · 37.7% Gemini 3.8 Live Extended Thinking · $3.50 · 68.6% Grok Voice Think Fast 2.0 · $4.80 · 56.5% GPT-Live-1 Astra · $5.83 · 67.9% Filled dots: the two Gemini 3.8 models. Settings as labelled by Google; its PDF says τ-Voice runs used thinking high.
The name on the headline number and the name in Search Live are the same, but they are two models. The one priced for everyone completes 30.1% of τ-Voice tasks; the one carrying the 68.6% costs more than four times as much per hour.

The base model does win something. Google says it took second place in the Speech Agent Arena, where people compare two hidden models in live conversations, and the leaderboard confirms it: Elo 1083, just behind Gemini 3.1 Flash Live (Minimal) at 1096. The arena prints its own 95% intervals, ±30 and ±27, and gives the base model a rank range of first to fourth. Its task success rate there is 93.2%. That sounds like a different model from the one that completes 30.1% of τ-Voice tasks, until you look at the arena’s scenarios: booking a dental appointment and ordering takeout. τ-Voice, as Sierra designed it, asks an agent to do things like change a flight under airline policy, while a simulated caller who can interrupt speaks over audio mixed with noise and telephone compression. Both numbers can be true at once. The Extended Thinking model, for what it’s worth, sits tenth in the arena at 990.

How wide a lead is

Here are the three tables Google cites, drawn on one scale.

Three leaderboards, and how far first place sits from second Three dot strips on a shared 0 to 100 percent scale. Artificial Analysis Speech to Speech Index: Gemini 3.8 Live Extended Thinking 82.6, GPT-Live-1 Astra 81.5, Grok Voice Think Fast 2.0 81.3, Gemini 3.8 Live 76.0, Gemini 3.1 Flash Live High 71.5 and Minimal 63.9; the lead is 1.1 points. τ-Voice, 278 tasks: Extended Thinking 68.6 as a mean of three trials, GPT-Live-1 Astra 67.9 from one trial, Grok 56.5, Gemini 3.1 Flash Live High 37.7, Gemini 3.8 Live 30.1, Gemini 3.1 Flash Live Minimal 26.2; the lead is 0.7 points, and a whisker marks plus or minus one standard error of a single pass over 278 tasks, about 2.8 points. τ³-Banking, 97 tasks: Extended Thinking 35.1, GPT-Live-1 Astra 32.0, xAI Realtime 16.5, Gemini 3.1 Flash Live High 11.3, GPT-Realtime 2 High 10.3; the base Gemini 3.8 Live is not reported; the lead is 3.1 points, and the single-pass standard error is about 4.7 points. THREE TABLES GOOGLE CITES · SHARED 0–100 SCALE · SCORES AS CHARTED 15 SEPT 2026 Gemini 3.8 Live Extended Thinking GPT-Live-1 Astra (no. 2) Gemini 3.8 Live Other models 0255075100 Speech to Speech Index Gap to no. 2: 1.1 · composite score Gemini 3.1 Flash Live (Minimal): 63.9 — Artificial Analysis via Google, 15 Sept 2026Gemini 3.1 Flash Live (High): 71.5 — Artificial Analysis via Google, 15 Sept 2026Grok Voice Think Fast 2.0 (High): 81.3 — Artificial Analysis via Google, 15 Sept 2026 Gemini 3.8 Live: 76.0 — Artificial Analysis via Google, 15 Sept 2026 GPT-Live-1 Astra (Medium): 81.5 — Artificial Analysis via Google, 15 Sept 2026 Gemini 3.8 Live Extended Thinking (High): 82.6 — Artificial Analysis via Google, 15 Sept 2026 82.681.576.0 τ-Voice · 278 tasks Gap: 0.7 · no. 2 scored on 1 trial Gemini 3.1 Flash Live (Minimal): 26.2% — Artificial Analysis via Google, 15 Sept 2026Gemini 3.1 Flash Live (High): 37.7% — Artificial Analysis via Google, 15 Sept 2026Grok Voice Think Fast 2.0 (High): 56.5% — Artificial Analysis via Google, 15 Sept 2026 Gemini 3.8 Live: 30.1% — Artificial Analysis via Google, 15 Sept 2026 ±1 standard error of one pass over 278 tasks at 67.9%, about 2.8 points — author's arithmetic GPT-Live-1 Astra (Medium): 67.9%, one trial — Artificial Analysis, Sept 2026 Gemini 3.8 Live Extended Thinking (High): 68.6%, mean of three trials — Artificial Analysis via Google, 15 Sept 2026 68.667.930.1 τ³-Banking · 97 tasks Gap: 3.1 · 3.8 Live not reported GPT-Realtime 2 (High): 10.3% — Google evaluation document, 15 Sept 2026Gemini 3.1 Flash Live (High): 11.3% — Google evaluation document, 15 Sept 2026xAI Realtime: 16.5% — Google evaluation document, 15 Sept 2026 ±1 standard error of one pass over 97 tasks at 32.0%, about 4.7 points — author's arithmetic GPT-Live-1 Astra (Medium): 32.0% — Google evaluation document, 15 Sept 2026 Gemini 3.8 Live Extended Thinking (High): 35.1% — Google evaluation document, 15 Sept 2026 35.132.0 Whiskers: ±1 standard error of a single pass over the task set around no. 2, assuming independent tasks. My arithmetic, not the evaluators'. The index is a composite of four components and has no such error bar.
On the two agent benchmarks, the gap between first and second is smaller than the sampling error of a single run over the task set. On τ-Voice the two dots overlap.

The τ-Voice gap is 0.7 points. Artificial Analysis, which runs the test, is open about how it gets its numbers, and its methodology page says each score is the mean of three independent trials over 278 scenarios: 50 airline, 114 retail and 114 telecom. A note on the leaderboard adds that GPT-Live-1 Astra, the runner-up, was scored on one trial. So the headline lead compares a three-run average against a single run. If the tasks are treated as independent, one pass over 278 of them at about 68% has a standard error near 2.8 points (my arithmetic). The lead is a quarter of that. On these numbers the two models are tied.

The Speech to Speech Index lead is 1.1 points, and the index needs care anyway. Google’s post calls it a “Quality Index”; Artificial Analysis, which built it, calls it the Speech to Speech Index. It has also changed shape twice since it launched in June. Version 2.0, dated August, weights four components equally: Big Bench Audio, τ-Voice, arena preference and arena task success. The Extended Thinking model ranks tenth on arena preference, so that component isn’t what puts it first.

The banking number is less settled still. Google’s blog calls the benchmark τ-Voice-banking. Its own PDF titles the chart “τ³-Banking Leaderboard” and describes a test in which agents search a large unstructured knowledge base and chain tool calls through banking workflows. It is the banking domain Sierra introduced as τ-knowledge in 2026: 698 documents, about 195,000 tokens, and tasks that need 9.5 tool calls on average. Artificial Analysis’ version runs 97 of those tasks and reports pass@1 “averaged across repeats,” without giving the number of repeats. The Gemini row doesn’t appear in the leaderboard text I pulled, so for Google’s 35.1% against Astra’s 32.0% I have only Google’s chart. A single pass over 97 tasks carries a standard error of nearly five points.

The metric the benchmark was built for

Google’s evaluation document settles one question in its first line of methodology: “All Gemini scores are pass @1 except where otherwise noted.” No exception is noted for the agent benchmarks. Pass@1 is the average chance that one attempt succeeds. It is what nearly every leaderboard reports, and for a voice agent it measures the wrong thing, which the benchmark’s own designers pointed out.

When Sierra introduced τ-bench in June 2024 (Yao et al.), the paper proposed a companion metric, passk: the chance that an agent succeeds on all k independent tries at the same task, averaged over tasks. The reasoning was that a customer-service agent that solves a problem sometimes is not an agent you can deploy. In that paper GPT-4o completed about 61% of retail tasks on one try, and its pass8 fell to about 25%. τ-Voice inherits the same tasks, tools and evaluator. Its launch post from 1 May says they are byte-for-byte identical to the text version’s.

Passk can’t be recovered from pass@1 alone, but it can be bounded. If every task is equally hard and tries are independent, passk is the single-attempt rate raised to the kth power. If every task is all-or-nothing, where the agent always solves it or never does, passk equals pass@1. Real benchmarks fall between those bounds. Artificial Analysis already ran three τ-Voice trials of the Extended Thinking model, so the data for pass3 exists. It isn’t published, and on 68.6% the bounds allow anything from 32.3% to 68.6%.

The published pairs show where measured agents land.

Single-attempt success against success on every attempt A dumbbell chart on a 0 to 100 percent scale. Three measured pairs: GPT-4.1 with a ReAct agent on AppWorld, from IBM Research, 77.4% average success across five runs against 53.0% of tasks solved in all five. GPT-5.5 at xhigh reasoning on Sierra's τ-Banking, 37.4% pass-one against 20.6% pass-four. GPT-5.2 at high reasoning on the same benchmark at launch, 25.5% against 9.3%. Two unreported pairs for Gemini 3.8 Live Extended Thinking, drawn as dashed ranges: on τ-Voice, 68.6% single-attempt, with a possible pass-three anywhere from 32.3% to 68.6%; on τ³-Banking, 35.1% single-attempt, with a possible pass-four anywhere from 1.5% to 35.1%. SINGLE-ATTEMPT RATE VS SUCCESS ON ALL K ATTEMPTS · PUBLISHED PAIRS AND GEMINI'S UNREPORTED RANGE Average single-attempt success Measured: solved on all k attempts Not reported: possible range 0%25%50%75%100% GPT-4.1 · AppWorld, 168 tasks IBM Research · k = 5 GPT-4.1 ReAct on AppWorld: 77.4% average over five runs — IBM Research, 15 Sept 2026 GPT-4.1 ReAct on AppWorld: 53.0% of tasks solved in all five runs — IBM Research, 15 Sept 2026 77.453.0 GPT-5.5 xhigh · τ-Banking Sierra · k = 4 GPT-5.5 xhigh on τ-Banking: 37.4% pass^1 — Sierra, 13 May 2026 GPT-5.5 xhigh on τ-Banking: 20.6% pass^4 — Sierra, 13 May 2026 37.420.6 GPT-5.2 high · τ-Banking Sierra, at launch · k = 4 GPT-5.2 high on τ-Banking at launch: 25.5% pass^1 — Sierra, 13 May 2026 GPT-5.2 high on τ-Banking at launch: 9.3% pass^4 — Sierra, 13 May 2026 25.59.3 Gemini 3.8 Live Ext. Thinking τ-Voice · 3 trials run · k = 3 Possible pass^3 range given 68.6% single-attempt: 32.3% to 68.6% — author's arithmetic Gemini 3.8 Live Extended Thinking on τ-Voice: 68.6% — Artificial Analysis via Google, 15 Sept 2026 68.6floor 32.3 Gemini 3.8 Live Ext. Thinking τ³-Banking · repeats not stated · k = 4 Possible pass^4 range given 35.1% single-attempt: 1.5% to 35.1% — author's arithmetic Gemini 3.8 Live Extended Thinking on τ³-Banking: 35.1% — Google evaluation document, 15 Sept 2026 35.1floor 1.5 Range: floor is the single-attempt rate to the kth power (every task equally hard, independent tries); ceiling is the rate itself (every task all-or-nothing). My arithmetic. Different agents, harnesses and task sets.
Every published pair loses a large share of its single-attempt score when the agent has to succeed every time. For Gemini's two agent scores, nobody has published the second number.

IBM Research published the cleanest recent example on the same day as Google’s launch. A ReAct agent on GPT-4.1, run five times at temperature zero over 168 AppWorld tasks, succeeded on 77.4% of runs and on all five runs for only 53.0% of tasks. The underlying paper calls that 24-point shortfall the consistency gap. IBM traces it to the serving side more than to sampling: on a hosted endpoint the probabilities drift slightly between calls, so a decision that was close to a tie can go the other way on the next run, and a multi-step task gives it many chances to. IBM’s guidelines narrowed the gap to 12 points, 81.0% against 69.0%.

Sierra’s own τ-knowledge post gives the banking-specific version. When the domain launched in March, GPT-5.2 at high reasoning passed 25.5% of tasks on the first try and 9.3% on all four. By May, GPT-5.5 at its top reasoning setting reached 37.4% and 20.6%. Those models, harnesses and task sets differ from Google’s, so none of this is an estimate for Gemini. As an illustration only: if Gemini’s 35.1% kept the same share GPT-5.5 kept, its pass4 would be about 19%, roughly one banking task in five handled correctly four times running.

The redial

In September 1964 the Bell System Technical Journal gave an entire issue to the No. 1 Electronic Switching System, the first large stored-program telephone exchange. Two of its papers deal with dependability, and they invert the usual computer priorities. In a computing centre, the overview paper argues, a stopped machine is a nuisance, because the job can be rerun, and a wrong answer is the disaster. In a telephone office it runs the other way: a total failure is the disaster, and a mishandled call is a nuisance to the customer, who has to redial. The maintenance-plan paper sets the objective that followed: no more than two hours of total downtime in the office’s 40-year design life.

A mishandled call could be written off as a nuisance for two reasons. The customer usually knew the call had failed, and the redial was close to an independent second draw from a very reliable machine. Retrying until something works is the world pass@k describes. The τ-bench paper says as much: pass@k suits tasks where a result can be checked, and customer service is not one of them.

A voice agent breaks the first condition more often than the second. τ-Voice scores the final state of the database, not the conversation, and Sierra’s launch post names the risk it was built to catch: agents that hold a pleasant conversation while quietly failing the task. A caller whose refund was filed against the wrong order, or not filed at all, has no reason to call back until the statement arrives. Nothing prompts a redial, so what the caller experiences is the first attempt and every attempt after it, which is passk.

That’s why the missing number matters more for this release than the 0.7 points. Google’s post says the models give developers and enterprises “the building blocks for reliable, production-ready voice agents.” Reliability has a measure in the benchmark family Google chose to cite, and Artificial Analysis appears to have enough τ-Voice trials in hand to compute part of it. When I wrote about a chess honeypot yesterday, the argument turned on denominators nobody had published. This is the same problem pointed at deployment. A 68.6% can mean an agent that solves two-thirds of problems every time, or one that solves most problems about two-thirds of the time. Those are very different products, and the published number cannot tell them apart. The harness question from Astra’s two ARC-AGI-3 scores applies here too: Google’s PDF says ServiceNow ran its EVA-Bench numbers through Google’s enterprise agent platform, while the τ-Voice and banking runs went through the public API.

None of this needs new experiments. Artificial Analysis could publish pass3 from the τ-Voice trials it has already run and score the runner-up on three trials like everyone else. Google could put the model’s full name next to every number.

References

  1. Ouyang, T., and Jaganathan, M. (2026, September 15). Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking. The Keyword, Google.
  2. Google DeepMind. (2026, September). Gemini 3.8 Audio (Live, Live Extended Thinking): Model evaluation: approach, methodology and results. PDF; results “as of September 2026.”
  3. Google DeepMind. (2026, September 15). Gemini 3.8 Audio (Live, Live Extended Thinking) model card.
  4. Artificial Analysis. (2026, accessed September 15). Speech to Speech leaderboard. Source for cost per hour, the τ-Voice trial-count note, and the index components.
  5. Artificial Analysis. (2026, accessed September 15). Speech to Speech benchmarking methodology, including the index version history.
  6. Artificial Analysis. (2026, accessed September 15). Speech Agent Arena leaderboard.
  7. Artificial Analysis. (2026, accessed September 15). τ³-Banking evaluation.
  8. Ray, S., Dhandhania, K., and Barres, V. (2026, May 1). τ-voice: benchmarking real-time voice agents on real-world tasks. Sierra.
  9. Shi, B., Zytek, O., Razavi, P., and Barres, V. (2026, May 13). τ-knowledge: benchmarking agents on real-world knowledge. Sierra.
  10. Duesterwald, E., et al. (2026, September 15). Your Agent Aced the Task. Will It Do It Again? IBM Research, Hugging Face blog.
  11. Duesterwald, E., Elder, B., Ngweta, L., Ubaru, S., and Zimon, M. (2026). Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course. arXiv:2609.08832, 8 September 2026.
  12. Yao, S., Shinn, N., Razavi, P., and Narasimhan, K. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv:2406.12045, 17 June 2024. Historical background.
  13. Keister, W., Ketchledge, R. W., and Vaughan, H. E. (1964). No. 1 ESS: System Organization and Objectives. Bell System Technical Journal, 43(5), 1831–1844, September 1964.
  14. Downing, R. W., Nowak, J. S., and Tuomenoksa, L. S. (1964). No. 1 ESS Maintenance Plan. Bell System Technical Journal, 43(5), 1961–2019, September 1964.

Standard errors, the passk bounds and the 19% illustration are my own calculations from the published figures.