OpenAI Launches GPT-Live-1, a Full-Duplex Voice Model That Delegates Deep Reasoning to GPT-5.5
OpenAI's new full-duplex voice models listen and speak at once and hand complex queries to GPT-5.5, replacing Advanced Voice Mode in ChatGPT.
Editor's Note ·
- Correction:
- The article states GPT-Live's launch 'comes two days before OpenAI publicly launched GPT-5.6.' GPT-Live launched on July 8, 2026, and GPT-5.6 reached general availability on July 9, 2026 -- a one-day gap, not two.
Overview
OpenAI has released GPT-Live-1 and GPT-Live-1 mini, a pair of full-duplex voice models the company describes as conversational models that “sound more natural and can handle turn-taking better,” according to TechCrunch. They are “full-duplex models, meaning they can speak and listen at the same time,” as TechCrunch reported. GPT-Live-1 mini now replaces Advanced Voice Mode as ChatGPT’s default voice experience, while the larger GPT-Live-1 is available to users on paid tiers, per TechCrunch.
What We Know
GPT-Live-1 is what OpenAI calls “a full-duplex audio language model — it processes incoming speech and generates outgoing speech concurrently, rather than waiting for a user to finish talking before formulating a response,” according to MLQ News. OpenAI’s own system card describes the models similarly, saying they “can listen and respond continuously instead of waiting for a clearly defined turn to end.”
The model “make[s] interaction decisions many times per second” about whether to “speak, continue listening, pause, interrupt, or invoke a tool,” according to MarkTechPost, a departure from earlier systems that relied on “silence-based” turn detection. During conversation, the system can produce short backchannel cues like “mhmm” or “yeah” while a user is speaking, MarkTechPost reported, and the model can also “stay silent for a long time and absorb the context of the conversation until it’s called upon,” per TechCrunch.
For harder queries, GPT-Live hands off work rather than trying to reason through it natively. “For queries requiring web search, deeper reasoning, or multi-step agent work, GPT-Live-1 delegates to GPT-5.5 running in the background, returning results into the live conversation without breaking flow,” MLQ News reported. Users can choose among three reasoning levels — Instant, Medium, and High — corresponding to different GPT-5.5 configurations: Instant uses GPT-5.5 Instant, Medium uses GPT-5.5 Thinking at medium effort, and High uses GPT-5.5 Thinking at high effort, according to MarkTechPost.
That delegation shows up in benchmark results. “GPT-Live-1 at its highest reasoning setting scored 84.2% on the GPQA scientific reasoning test, up from 45.3% for Advanced Voice Mode,” according to MLQ News. On BrowseComp, a test of agentic web search, “GPT-Live-1 scored 75.2% compared to 0.7% previously,” MLQ News reported. In head-to-head human evaluations, “participants chose GPT-Live-1 over Advanced Voice Mode in 75.7% of comparisons,” per MLQ News, and MarkTechPost separately reported that in head-to-head tests, both new models were “strongly preferred over Advanced Voice Mode” across categories measuring overall preference, turn-taking, interruptions, and flow, according to MarkTechPost. The models also outperformed on a telecom-specific benchmark, τ³-Voice Telecom, which tests multi-turn support tasks, per MarkTechPost.
On availability, GPT-Live-1 rolls out to subscribers on OpenAI’s Go, Plus, and Pro tiers, while GPT-Live-1 mini — described as “a lighter version” — is available to free-tier users, according to MLQ News. Both models “roll out to ChatGPT users globally today,” MarkTechPost reported. The launch also includes nine remastered voices and visual cards that can display information such as weather, stocks, and sports scores, alongside continued support for search, memory, images, and file uploads, per MarkTechPost.
Atty Eleti, ChatGPT Voice’s product lead, framed the release as a step toward a broader shift in how people interact with AI. “Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing,” Eleti said, according to TechCrunch.
On safety, OpenAI’s own system card says the new models achieved “equal or better safety performance” than Advanced Voice Mode across most categories tested, using both real audio examples voluntarily shared by users and synthetically generated edge-case prompts. The card does note two small exceptions: GPT-Live-1’s score for emotional-reliance handling slipped from 0.88 to 0.82, and GPT-Live-1 mini’s score for sexual-content handling slipped from 0.97 to 0.95. OpenAI also said neither model “could plausibly be considered High in any of” its Preparedness Framework’s tracked risk categories, which cover biological and chemical risk, AI self-improvement, and cybersecurity, according to the system card. At runtime, the system checks inputs and outputs continuously and, per the system card, “when potentially unsafe content is detected, the system can steer or interrupt the response, play a spoken safety message, provide support resources in text, or, in higher-risk cases, end the voice conversation.” The models also include safeguards intended to give age-appropriate responses to teenagers and to surface support resources if a conversation turns to topics like self-harm, TechCrunch reported.
The launch comes two days before OpenAI publicly launched GPT-5.6, the text model whose predecessor, GPT-5.5, now handles GPT-Live’s delegated reasoning and search tasks. It also follows Google’s April release of Gemini 3.1 Flash Live, a real-time voice model that similarly targets natural, low-latency conversation — underscoring that full-duplex, agentic voice interfaces have become a contested front among frontier AI labs.
What We Don’t Know
OpenAI has not set a date for opening GPT-Live to its API. MarkTechPost reported that “the API is planned soon” but is not yet available, according to MarkTechPost, while MLQ News reported API access is coming but that OpenAI did not provide a specific date. Video, screen sharing, and full multilingual parity are not available at launch, per MarkTechPost, and it remains unclear when those capabilities will arrive. TechCrunch also noted limitations in a live-translation demo, where Hindi output carried a “heavy American accent” and sounded “unnatural,” according to TechCrunch, suggesting multilingual performance still has room to improve even within supported languages.