The Race to Sub-Second Voice AI: Speech-to-Speech vs. the Old Pipeline
Human conversation has a rhythm: the typical gap between one person finishing a sentence and the other starting theirs is about 200 milliseconds. For years, voice AI couldn't get close to that. A wave of 2025-2026 architecture changes is closing the gap -- here's what actually changed, sourced to the technical publications and independent benchmarks, not just vendor marketing copy.
The numbers behind the race
~200ms
The typical human-to-human conversational turn-taking gap, per peer-reviewed research published in PNAS -- the benchmark every voice AI platform is racing toward.
Single-model shift
OpenAI's gpt-realtime, Kyutai's Moshi and Hume AI's EVI 3 all replaced the old speech-to-text-to-speech pipeline with one native audio model.
0.44s fastest
Time-to-first-audio for the fastest model on the independent Artificial Analysis Speech-to-Speech leaderboard, among 34+ models tested.
No system wins on all axes
Sesame's TurnBench benchmark found no tested platform is simultaneously fast, selective and high-recall at handling real conversational turn-taking.
The architecture shift: from three models to one
For most of the last decade, voice AI worked as a relay race: speech-to-text converted the caller's audio to words, a language model decided what to say back, and text-to-speech converted that answer into audio -- three separate models, each adding its own processing delay before the next stage could even start.
OpenAI's gpt-realtime, launched August 28, 2025, is explicit about breaking that pattern: "unlike traditional pipelines that chain together multiple models across speech-to-text and text-to-speech, the Realtime API processes and generates audio directly through a single model." Kyutai's research model Moshi takes a similar full-duplex approach -- it models the user's and system's audio as parallel streams through a neural audio codec plus a transformer architecture, with a published theoretical latency of 160ms and roughly 200ms in practice, which the paper's authors position directly against the ~230ms average human turn-taking gap they cite. Hume AI's EVI 3, launched May 29, 2025, is described as a unified speech-language model rather than a cascade, with a vendor-stated claim of "under 300ms on state-of-the-art hardware" -- a figure worth treating as a vendor claim rather than an independently verified one.
Sources: OpenAI, "Introducing gpt-realtime," Aug 28, 2025; Kyutai, Moshi paper (PDF); Hume AI, "Introducing EVI 3," May 29, 2025.
The text-to-speech side of the race
Even where the STT-LLM-TTS pipeline hasn't been fully replaced, individual components have gotten dramatically faster. Cartesia's Sonic, launched May 31, 2024, uses a state-space model rather than a transformer specifically for speech generation, with an original claimed model latency of 135ms. By August 2026, Cartesia's Sonic-3.6 update claimed sub-90ms time-to-first-audio and topped both of the independent Artificial Analysis speech arenas -- the provider Elo leaderboard (1,283) and the "controlled voice" board (1,123), ahead of its own prior version and ElevenLabs' Eleven v3.
It's worth being precise about what these numbers measure: a "time-to-first-audio" or "model latency" figure for a TTS component is not the same as the full round-trip time a caller experiences on an actual phone call, which also includes speech recognition, network transit, and the language model's own "thinking" time. A Telnyx comparison piece explicitly warns against comparing a vendor's component-only latency claim (such as a TTS-only figure) to genuine voice-to-voice round-trip measurements -- exactly the kind of category error that makes vendor latency marketing hard to compare at face value.
Sources: Cartesia, "Announcing Sonic," May 31, 2024; MarkTechPost, Aug 18, 2026; Artificial Analysis, Speech-to-Speech leaderboard; Telnyx, voice AI latency comparison.
What independent benchmarks actually show
Two sources stand out for measuring across vendors rather than reporting one company's own numbers. The Artificial Analysis Speech-to-Speech leaderboard tests 34+ models on time-to-first-audio: as of this writing, Deepslate Opal leads at 0.44 seconds, Gemini 2.5 Flash Native Audio Dialog follows at 0.63 seconds and Grok Voice Think Fast 2.0 High at 0.70 seconds -- with slower models trailing at 4-5+ seconds, a meaningful spread for a category where sub-second response is the whole point.
A Telnyx-aggregated third-party test by Tested Media measured 500 production calls per platform and found Vapi at a 720ms median, Retell at 680ms, and Bland at 850ms -- genuine voice-to-voice figures rather than component claims. Sesame's own TurnBench, published August 28, 2026, goes further by measuring end-of-turn and interruption-handling latency specifically: OpenAI's Realtime API scored 282ms end-of-turn with Server VAD (though 793ms with its more accurate Semantic VAD setting), Moshi scored 702ms, and Gemini 3.1 Live scored 1,234ms. Sesame's own conclusion, despite being the benchmark's vendor-author: "no system is fast, selective, and high-recall at the same time" -- every platform tested trades off one property against another rather than winning outright.
Sources: Artificial Analysis, Speech-to-Speech leaderboard; Telnyx, voice AI latency comparison; Sesame, "TurnBench," Aug 28, 2026; Stivers et al., PNAS, 2009.
Frequently asked questions
What is speech-to-speech AI, and how is it different from older voice AI?
Older voice AI chains three separate models: speech-to-text, then a language model, then text-to-speech. Speech-to-speech (native audio) models process and generate audio directly through a single model, cutting out the hand-offs between stages. OpenAI's gpt-realtime, Kyutai's Moshi and Hume AI's EVI 3 all use this native approach.
How fast is human conversational turn-taking, and why does that matter for voice AI?
Research published in the Proceedings of the National Academy of Sciences found the typical gap between conversational turns across languages is only about 200 milliseconds. Voice AI platforms use that figure as the target to beat -- a system that takes much longer to respond feels noticeably less natural than a real conversation.
Should I trust a voice AI vendor's own latency claims?
Treat them carefully. Some vendor-quoted figures measure only one component (e.g. text-to-speech generation time) rather than the full round trip a caller actually experiences. Independent, third-party voice-to-voice tests -- like production-call measurements from Tested Media covering Vapi, Retell and Bland -- are more representative than a single vendor's own marketing number.
Kamaljeet Singh Sidhu
Founder & CEO of Botnira and CEO of The DigiSparrow. Reviews the independent industry research published on this hub.
Read full bio →Related reading
Real-Time AI Voice Translation: What's Actually Possible in 2026
The technical and market state of multilingual voice AI, sourced to product launches and one peer-reviewed clinical study.
Read the report →The AI Receptionist: How It Works, and What the Data Shows
How response latency, language coverage and hallucination-avoidance are measured in a real deployment.
Read the report →Agentic AI in Contact Centers: The 2029 Prediction
Gartner's real, verified prediction on how much of customer service agentic AI will resolve autonomously.
Read the report →