OpenAI Launches GPT-Live-1, a Full-Duplex Voice Model for AI Phone Agents
OpenAI released GPT-Live-1 in its API on September 10, 2026 -- a voice model built to listen and speak at the same time, rather than waiting for a caller to finish before responding. Priced at $0.05 a minute, it targets the exact weak point that has dogged AI phone agents since the category began: the awkward, laggy back-and-forth that gives away that a caller is talking to a machine.
What OpenAI shipped on September 10
OpenAI made GPT-Live-1 available to developers through its API on September 10, 2026, exposing the full-duplex voice model that already powers ChatGPT's voice mode to anyone building a voice agent. The model is served through a dedicated Live endpoint (v1/live/sessions) rather than the older Chat Completions, Assistants or Realtime endpoints, and OpenAI's own model documentation lists a knowledge cutoff of July 31, 2025, support for both audio and text as input and output, streaming, and function calling, with concurrent-session rate limits that scale from 25 sessions on the lowest usage tier up to 500 on the highest.
The core idea is architectural, not cosmetic. Almost every AI phone agent shipped to date -- including most of the products this section has covered from Intermedia to Gupshup to SkySwitch -- is built as a cascaded pipeline: a speech-to-text model transcribes the caller, a separate reasoning model decides what to say, and a text-to-speech model speaks the reply, with each handoff adding latency and a chance to lose conversational context. GPT-Live-1 instead reasons over incoming and outgoing audio in one continuous stream, so the model can register an "mm-hm," a caller talking over it, or a mid-sentence change of direction as it happens, without stopping to wait its turn.
Source: OpenAI, "Introducing GPT-Live", September 10, 2026; OpenAI API documentation, GPT-Live-1 model page.
The numbers behind the "feels human" claim
OpenAI's benchmark comparisons put concrete figures behind the pitch. Measured against GPT-Realtime-2.1, its previous turn-based voice model, GPT-Live-1 cut turn-taking latency from 1.41 seconds down to 0.798 seconds, more than doubled its pass rate on OpenAI's own Tau3 Voice Intelligence benchmark (from 45.7% to 86.2%), and nearly doubled its interactivity score on Full Duplex Bench (from 45.4% to 80.10%). OpenAI also reports that, when paired with its GPT-6 Astra reasoning model, the combination showed roughly 80% fewer interruptions than prior turn-based systems generated -- interruptions being the moments where an AI agent talks over a caller or starts responding before the caller has actually finished a thought, one of the most common complaints about early-generation AI phone agents.
Pricing is $0.05 per minute for the voice layer itself, billed by the second rather than rounded up, with the backend reasoning model and any tool calls charged separately -- so total cost still depends on which model a developer pairs GPT-Live-1 with, similar to how OpenAI already prices its Realtime API. The model ships with 12 voice options spanning different accents and languages, native turn detection, keyword biasing and alphanumeric recognition for things like confirmation numbers or spelled-out names, and telephony support aimed squarely at phone-based use cases: reservations, order updates, customer support and scheduling calls.
Source: Unite.AI, "OpenAI's GPT-Live-1 Arrives in the API at $0.05 Per Minute", September 2026.
Who's already using it
OpenAI named several early adopters already running GPT-Live-1 in production. Yelp's Host product uses it for its restaurant-facing phone agent; Yelp CTO Alex Levy said callers are "speaking fuller, more natural sentences," which the company takes as a signal that "the experience on the other end feels genuinely different." Language-learning app Speak, per its CTO Andrew Hsu, cut interruptions during learner pauses by nearly 80% compared with its prior system -- a meaningful detail for a product where the whole point is letting a nervous learner think mid-sentence without the AI barging in. Intercom's AI support agent Fin is using the model to move voice support "toward the natural flow of a phone call," according to Fin COO Jordan Neil, while Cognition, maker of the AI coding agent Devin, is using GPT-Live-1 for live voice collaboration between a developer and the agent. One unnamed customer reportedly eliminated roughly 23,000 lines of custom real-time conversation code -- an 80% reduction in the codebase previously needed to stitch together a cascaded voice pipeline -- once it moved onto the unified model.
Source: Unite.AI, "OpenAI's GPT-Live-1 Arrives in the API at $0.05 Per Minute", September 2026.
What this means for the AI voice and calling industry
The cascaded voice pipeline is losing ground
Most AI phone agents on the market today are still stitched together from separate speech-to-text, reasoning and text-to-speech components. A foundation model vendor shipping a genuinely unified alternative, at API scale and developer pricing, puts pressure on every voice AI platform still built the old way.
Latency and interruptions are now the benchmark, not accuracy
OpenAI is competing on turn-taking latency and interruption rates rather than raw transcription quality -- a sign that basic speech recognition is now table stakes, and the real battleground for AI phone agents has shifted to how naturally the conversation actually flows.
Infrastructure vendors, not just app builders, will feel this
Voice AI infrastructure companies that sell orchestration around cascaded pipelines -- turn detection, interruption handling, latency tuning -- now have to justify that layer against a single model that claims to do it natively, which could reshape where margin sits in the voice AI stack.
Related reading
Microsoft Launches MAI-Transcribe-2, Undercutting Meta and OpenAI on Price
Another foundation model vendor competing on price and benchmark performance in the audio AI space, this time on speech-to-text rather than full-duplex conversation.
Read the article →Meta Launches Muse Voice Transcribe, Undercutting Rival Speech AI on Price
Meta's own real-time audio perception model launch, part of the same wave of large AI labs moving directly into voice and speech infrastructure.
Read the article →AI Receptionist Research Hub
How foundation-model voice layers like GPT-Live-1 compare to purpose-built small-business AI receptionists like Botnira on reliability, setup and cost.
Browse the research →