OpenAI launches GPT-Live: talk to AI without taking turns — hard questions auto-route to GPT-5.5
- OpenAI has released a new-generation voice model, GPT-Live, built on a full-duplex architecture that keeps processing input while it's still generating speech — making natural, listen-while-speaking conversation possible.
- For tasks needing web search, deep reasoning, or complex work, GPT-Live hands the question off to GPT-5.5 running in the background in real time, without interrupting your side of the conversation.
- Two versions — GPT-Live-1 and GPT-Live-1 mini — roll out today to ChatGPT users worldwide: Go / Plus / Pro default to GPT-Live-1, while Free users default to the mini version.
- In human head-to-head evaluations, both GPT-Live-1 and mini were clearly preferred over Advanced Voice Mode; they also beat it on GPQA, BrowseComp, and τ³-Voice Telecom.
- Voice mode adds visual cards for weather, stocks, sports, and more; reasoning effort is now adjustable (Instant / Medium / High); and OpenAI has strengthened safety protections for self-harm and emotional dependence in voice specifically.
OpenAI launches GPT-Live: voice conversation no longer takes turns
On July 8, 2026, OpenAI released a new-generation voice model, GPT-Live, and made it the new underlying model for ChatGPT Voice, rolling out to users worldwide today.
What GPT-Live does is let you talk to AI without taking turns: it can listen to you and respond at the same time, and you can interrupt or jump in with a follow-up mid-sentence. When a question needs a web search or real thinking, it quietly hands the work to the stronger GPT-5.5 running in the background — and your side of the conversation never drops.
Why it matters: The previous Advanced Voice Mode had to wait until you finished speaking before it could respond. GPT-Live achieves listen-while-speaking through a full-duplex architecture, and it beats Advanced Voice Mode on GPQA, BrowseComp, and τ³-Voice Telecom. GPT-Live-1 and GPT-Live-1 mini are both rolling out globally starting today.
Two waveforms can rise and fall at the same time without interrupting each other — that's what full-duplex conversation looks like.
Where the old approach breaks down: relay loses information, turn-taking gets talked over
To see what's new about GPT-Live, you first need to see where the previous two generations of voice tech got stuck. Both let you converse with AI, but each has its own awkward flaw. Laying the three architectures side by side makes it clearest.
(cascaded system)
but still takes turns
and speaks↓ Silence from you
= "you're done"↓ Then it replies
speaking (full-duplex)
per second
It's like a relay race: speech-to-text, the AI thinking up an answer, and text-to-speech are the three runners. Every handoff is a chance to drop the baton — losing a bit of the information in what you said.
Here's the same line — "hey, got a minute to chat" — spoken by all three architectures. OpenAI provided real recordings, so you can compare them directly:
Breakthrough 1: one model listens and speaks at once, no need to wait your turn
GPT-Live's first change is a full-duplex architecture (meaning it can listen and speak at the same time) built specifically for ongoing conversation. It no longer splits a conversation into separate turns — it keeps processing your input and keeps generating a response, continuously and simultaneously.
Because listening and speaking run at once, the model can make many decisions every second: speak now, keep listening, pause, interrupt, or call a tool. The result is a more natural back-and-forth, better timing, and it can even do real-time translation.
A walkie-talkie only lets one person talk at a time; the other side has to wait. A phone call lets both sides jump in at once and catch what the other says. A full-duplex voice model switches AI from walkie-talkie mode to phone-call mode.
You speak → let go → it speaks → finishes → then it's your turn. Pause in the middle to think, and it's likely to read that as "done" and cut in.
While you're talking, it's already listening and judging. When you pause to think something through, it waits quietly; when you signal it to jump in, it speaks — and it'll use "mm-hm" or "got it" to let you know it's following along.
Breakthrough 2: it handles everyday chat itself, and hands off the hard problems to a stronger model
The second change is splitting "keeping you company" from "doing the heavy lifting." GPT-Live focuses on sustaining the conversation; when a question needs a web search, real reasoning, or a chain of agentic actions, it delegates that hard problem to another, stronger model — at launch, that's GPT-5.5This backend engine isn't fixed: whenever OpenAI ships a stronger frontier model, GPT-Live can just plug in the new one behind it — no need to redo the interaction layer.. While the delegation is happening, your conversation keeps going, and the result gets folded back in once it's ready.
Here's a concrete scenario to see how this plays out: you're driving and ask ChatGPT by voice, "Can you check if there are still tickets for tomorrow's Beijing-to-Shanghai flight?" That's a job that needs a web lookup. GPT-Live keeps chatting with you while handing the task to the backend — and once the backend has an answer, it gets folded into the conversation, so you never feel a stall.
Benchmark results: both chat feel and expert-level questions beat the old model
For this launch, OpenAI ran new human evaluations to measure how comfortable and smooth conversations feel, and also compared several general-capability benchmarks against Advanced Voice Mode. All results below are OpenAI's own reported figures.
How the human head-to-head evaluation worked: people had 5–10 minute conversations and rated them across five dimensions — overall preference, turn-taking, handling interruptions, fluency, and naturalness — indicating which conversation they preferred each time, GPT-Live or Advanced Voice Mode.
There's also a separate conversation-rating test scored independently (out of 7 points, covering conversational fluency and overall comfort):
| Model | Conversational fluency | Comfort |
|---|---|---|
| GPT-Live-1 | 4.96 | 5.19 |
| GPT-Live-1 mini | 4.33 | 4.47 |
| Advanced Voice Mode | 3.80 | 3.82 |
These three general-purpose benchmarks weren't built for voice specifically — they test the model's underlying capability, again compared against Advanced Voice Mode:
τ³-Voice Telecom tracks both task-completion rate and time taken — higher reasoning effort means more accurate but slower:
| Model | Median task time | Task success rate |
|---|---|---|
| AVM | 385.5 sec | 29.5% |
| GPT-Live-1 mini | 290.9 sec | 39.5% |
| GPT-Live-1 · Instant | 231 sec | 37.3% |
| GPT-Live-1 · Medium | 295 sec | 59.6% |
| GPT-Live-1 · High | 386 sec | 63.4% |
Open ChatGPT Voice — here's what's actually different now
More than 150 million people use ChatGPT's voice and dictation features every week — as a hands-free daily assistant, for practicing a language, for bedtime stories, or just to chat during a commute. Starting today, tapping the voice button means you're using GPT-Live.
More like talking to a real person, and better at listening too
You can jump in with a question at any time, pause to gather your thoughts, or ask it to slow down. It'll use responses like "mm-hm" or "got it" to show it's following along. All nine voices have been remastered as well.
Even a brief pause to find the right words is likely to be read as "you're done," and it cuts in and interrupts you.
It waits quietly, without rushing to jump in. Tell it to stay quiet and just listen, and it will. In noisy settings — traffic, other people talking — it also stays better focused on your voice.
Smarter answers when you need them
Voice can now call on the latest frontier model, and you can choose the reasoning effort yourself: pick Instant for speed, or Medium/High when you want it to think harder.
Some answers are more useful when you can see them
While it's talking, ChatGPT can now show you cards directly — for weather, stocks, sports scores, and more — without needing to type anything to confirm. Voice still supports search, memory, and image/file uploads as before.
Two more official screenshots from OpenAI:
Safety design built specifically for voice
On top of the latest model's baseline safety, GPT-Live adds dedicated safety training and protections tailored to voice as a new medium.
Safety testing that's closer to real-world use
OpenAI extended its safety testing to audio-native evaluations, built a batch of tests using synthetic audio, and had internal experts run red-teaming focused specifically on voice-related risks. Per the original post, GPT-Live matched or beat Advanced Voice Mode across nearly every evaluation area. The key risk areas covered include:
Protections that can step in mid-speech
Because voice conversation happens in real time, OpenAI built protections that can act while the model is still speaking. If the system detects potentially unsafe content, it can do one of three things:
For conversations involving self-harm, OpenAI adapted ChatGPT's support pathways for the voice version, including expert-reviewed crisis-hotline support.
Teen protections and voice safeguards
For teen users, the model's training includes age-appropriate behavior. Parents can use Parental Controls to decide whether their teen can use ChatGPT Voice at all; in high-risk situations showing signs of self-harm or suicidal intent, a linked parent may be notified. In addition, GPT-Live only uses a fixed set of preset voices, built with protections in place and not designed to mimic real people's voices.
Can you use it now? Availability and current limitations
GPT-Live is rolling out to ChatGPT users worldwide starting now, covering iOS, Android, and ChatGPT.com. Here's a breakdown of which version each plan defaults to.
A few current limitations worth knowing: GPT-Live has been optimized for ChatGPT's most-used languages, and some languages may still come with a non-native accent or reduced fluency — OpenAI says it's working on improvements. At launch, it doesn't yet support voice combined with video or screen sharing, though those are coming soon. An API is also on the way; developers and businesses can sign up on a waitlist form to be notified.
- For everyday voice Q&A, language practice, or chatting during a commute, interruptions, thinking pauses, and requests to slow down are now recognized naturally — no more getting talked over or misread as often.
- For web lookups, expert-level science questions, or complex multi-step tasks, voice conversation can now reach accuracy close to a deep-reasoning model, since GPT-5.5 is handling it in real time behind the scenes.
- For weather, stocks, sports scores, and similar lookups, voice conversation now shows visual cards directly — no need to type to confirm.
- Parents can use Parental Controls to manage whether teens can use voice features, and high-risk conversations involving self-harm or suicidal intent will notify a linked parent.
GPT-Live is built on a full-duplex architecture, meaning it can listen and speak at the same time. For questions that need web search, deeper reasoning, or more complex work, it delegates to our latest frontier model behind the scenes, then brings the result back into the conversation once it's ready. OpenAI, "Introducing GPT-Live," July 8, 2026
Talking to AI by voice: from a "walkie-talkie that takes turns" to "a phone call that listens while it speaks"
OpenAI launches the GPT-Live voice model, which listens while speaking and auto-delegates hard questions to GPT-5.5 — the whole thing in one page with visuals.
↓ One page, full story · includes an animated figure
ChatGPT Voice lets you talk instead of type with AI — asking for directions on a commute, practicing a language, telling bedtime stories. 150 million people use it every week. It's always had one awkward flaw: only one side can talk at a time.
✘ Can't listen and speak at the same time, or let you interrupt at will, like a phone call
It relies on "you've gone quiet" to decide it's its turn. But if you just pause to find the right words, or there's a bit of background noise, it can read that as "you're done" and cut in.
GPT-Live runs on a new architecture that can listen and respond at the same time: when you pause, it waits quietly; when you interrupt, it stops. For questions that need a web search or real thinking, it hands the work to the stronger GPT-5.5 in the background, without breaking your side of the conversation.
You finish → let go → it speaks → finishes → then it's your turn. Pause in the middle and it's easy to get read as "done," and get talked over.
While you're talking, it's already listening and judging. When you pause to think, it waits quietly; when you signal it, it speaks — and uses "mm-hm" or "got it" to show it's following.
But "keeping you company" and "researching, reasoning" are two very different jobs — one model doing both alone risks dropping the ball on one of them. So how does GPT-Live keep both running smoothly? The figure below makes it clear.
Take an example: driving and asking by voice about flight availability. See how GPT-Live keeps chatting up front while handing the hard part to the backend.
The most tangible official number is a phone-support task: how much time does the same job take on the old version versus the new one?
how to phrase this...
you're done — cuts right in
take your time
interrupt and it stops
for tomorrow?
what time are you leaving?
it's real work
- × These numbers are all OpenAI's own tests
- × No third party has reproduced them yet
- × The faster the tier, the lower the success rate
