Product Launch · XiaoHu Explains

OpenAI launches GPT-Live: talk to AI without taking turns — hard questions auto-route to GPT-5.5

A true listen-while-speaking voice model goes live today, and it can hand complex tasks to GPT-5.5 in real time.
30-Second Overview
  • OpenAI has released a new-generation voice model, GPT-Live, built on a full-duplex architecture that keeps processing input while it's still generating speech — making natural, listen-while-speaking conversation possible.
  • For tasks needing web search, deep reasoning, or complex work, GPT-Live hands the question off to GPT-5.5 running in the background in real time, without interrupting your side of the conversation.
  • Two versions — GPT-Live-1 and GPT-Live-1 mini — roll out today to ChatGPT users worldwide: Go / Plus / Pro default to GPT-Live-1, while Free users default to the mini version.
  • In human head-to-head evaluations, both GPT-Live-1 and mini were clearly preferred over Advanced Voice Mode; they also beat it on GPQA, BrowseComp, and τ³-Voice Telecom.
  • Voice mode adds visual cards for weather, stocks, sports, and more; reasoning effort is now adjustable (Instant / Medium / High); and OpenAI has strengthened safety protections for self-harm and emotional dependence in voice specifically.
This article is an explainer based on OpenAI's official launch blog post. The benchmark results, preference data, and safety-testing conclusions cited here are OpenAI's own self-reported figures and internal testing — none have been independently verified by a third party.
1Launch

OpenAI launches GPT-Live: voice conversation no longer takes turns

On July 8, 2026, OpenAI released a new-generation voice model, GPT-Live, and made it the new underlying model for ChatGPT Voice, rolling out to users worldwide today.

What GPT-Live does is let you talk to AI without taking turns: it can listen to you and respond at the same time, and you can interrupt or jump in with a follow-up mid-sentence. When a question needs a web search or real thinking, it quietly hands the work to the stronger GPT-5.5 running in the background — and your side of the conversation never drops.

Why it matters: The previous Advanced Voice Mode had to wait until you finished speaking before it could respond. GPT-Live achieves listen-while-speaking through a full-duplex architecture, and it beats Advanced Voice Mode on GPQA, BrowseComp, and τ³-Voice Telecom. GPT-Live-1 and GPT-Live-1 mini are both rolling out globally starting today.

Official OpenAI demo video (Chinese/English captions): a real voice conversation between GPT-Live and a "grandma" character, showing off listen-while-speaking, being interrupted at any time, and waiting quietly — the core full-duplex traits. Source: OpenAI, "Introducing GPT-Live."
Your voiceAI's voice

Two waveforms can rise and fall at the same time without interrupting each other — that's what full-duplex conversation looks like.

2The Old Way

Where the old approach breaks down: relay loses information, turn-taking gets talked over

To see what's new about GPT-Live, you first need to see where the previous two generations of voice tech got stuck. Both let you converse with AI, but each has its own awkward flaw. Laying the three architectures side by side makes it clearest.

Generation 1
Three models in relay
(cascaded system)
You speak Speech to text AI thinks of an answer Text to speech
Each step has to wait for the last one to finish. Tone and inflection can get lost at each handoff, and replies come out slow and stiff.
Generation 2
One model handles both,
but still takes turns
One model listens
and speaks
Silence from you
= "you're done"
Then it replies
Faster and smoother, but only one side can talk at a time. Pause to think, or a bit of background noise, and it might read that as "you're done" and jump in and talk over you.
GPT-Live
Listen while
speaking (full-duplex)
Keeps listening Keeps responding Decides multiple times
per second
Listening and speaking happen at once, constantly judging whether to talk, listen, wait, or interrupt — closing the gap left by the previous two generations. The next two sections dig into how it pulls this off.
An Analogy · Cascaded systems

It's like a relay race: speech-to-text, the AI thinking up an answer, and text-to-speech are the three runners. Every handoff is a chance to drop the baton — losing a bit of the information in what you said.

Here's the same line — "hey, got a minute to chat" — spoken by all three architectures. OpenAI provided real recordings, so you can compare them directly:

Generation 1 · Cascaded system
Slow and stiff response, with long pauses. (Standard Voice Mode, using GPT-5.5 Instant)
Generation 2 · Turn-based (AVM)
Faster and smoother, but multi-turn exchanges still feel stiff. (ChatGPT Advanced Voice Mode)
GPT-Live · Full-duplex
Fast, natural, expressive response, and it listens more attentively too. (GPT-Live-1, using GPT-5.5 Instant)
3Breakthrough 1 · Full-duplex

Breakthrough 1: one model listens and speaks at once, no need to wait your turn

GPT-Live's first change is a full-duplex architecture (meaning it can listen and speak at the same time) built specifically for ongoing conversation. It no longer splits a conversation into separate turns — it keeps processing your input and keeps generating a response, continuously and simultaneously.

Core Innovation

Because listening and speaking run at once, the model can make many decisions every second: speak now, keep listening, pause, interrupt, or call a tool. The result is a more natural back-and-forth, better timing, and it can even do real-time translation.

Official OpenAI demo video (Chinese/English captions): an OpenAI staffer explains the "always-present" capability and demonstrates it live — brewing pour-over coffee while GPT-Live calls out the timing throughout, staying with the multi-step task and jumping in whenever needed. Source: OpenAI, "Introducing GPT-Live."
An Analogy · Full-duplex

A walkie-talkie only lets one person talk at a time; the other side has to wait. A phone call lets both sides jump in at once and catch what the other says. A full-duplex voice model switches AI from walkie-talkie mode to phone-call mode.

Walkie-talkie mode · Old version

You speak → let go → it speaks → finishes → then it's your turn. Pause in the middle to think, and it's likely to read that as "done" and cut in.

Phone-call mode · GPT-Live

While you're talking, it's already listening and judging. When you pause to think something through, it waits quietly; when you signal it to jump in, it speaks — and it'll use "mm-hm" or "got it" to let you know it's following along.

Official OpenAI demo video (Chinese/English captions): a real listen-while-speaking conversation with GPT-Live-1 (backed by GPT-5.5 Instant) — you can hear directly how it differs from the walkie-talkie mode above. Source: OpenAI, "Introducing GPT-Live."
Under the full-duplex architecture, the model makes these calls repeatedly, every second
Speak Keep listening Pause Interrupt Call a tool
4Breakthrough 2 · Delegation

Breakthrough 2: it handles everyday chat itself, and hands off the hard problems to a stronger model

The second change is splitting "keeping you company" from "doing the heavy lifting." GPT-Live focuses on sustaining the conversation; when a question needs a web search, real reasoning, or a chain of agentic actions, it delegates that hard problem to another, stronger model — at launch, that's GPT-5.5This backend engine isn't fixed: whenever OpenAI ships a stronger frontier model, GPT-Live can just plug in the new one behind it — no need to redo the interaction layer.. While the delegation is happening, your conversation keeps going, and the result gets folded back in once it's ready.

Core Innovation

Here's a concrete scenario to see how this plays out: you're driving and ask ChatGPT by voice, "Can you check if there are still tickets for tomorrow's Beijing-to-Shanghai flight?" That's a job that needs a web lookup. GPT-Live keeps chatting with you while handing the task to the backend — and once the backend has an answer, it gets folded into the conversation, so you never feel a stall.

Front-end · GPT-Live (conversation never stops) Back-end · GPT-5.5 (web search · reasoning) You ask "Are there tickets for tomorrow's flight?" GPT-Live keeps chatting "Let me check — what time are you leaving?" GPT-Live reports back "3 flights still have seats, earliest is 8 AM" conversation continues GPT-5.5 searches flight availability delegates task returns result
Conversation never stops: while the backend spends those few seconds looking up flights, the front-end GPT-Live doesn't freeze — it keeps confirming your departure time with you. Once the answer comes back, it slots naturally into the conversation. This split also means GPT-Live's backend engine can keep upgrading as new models arrive.
Real conversation recording · Delegating a deep task
GPT-Live delivers a fast, natural response while GPT-5.5 handles the search task in the background. (GPT-Live-1, using GPT-5.5 Instant) Source: OpenAI, "Introducing GPT-Live."
5Benchmarks

Benchmark results: both chat feel and expert-level questions beat the old model

For this launch, OpenAI ran new human evaluations to measure how comfortable and smooth conversations feel, and also compared several general-capability benchmarks against Advanced Voice Mode. All results below are OpenAI's own reported figures.

How the human head-to-head evaluation worked: people had 5–10 minute conversations and rated them across five dimensions — overall preference, turn-taking, handling interruptions, fluency, and naturalness — indicating which conversation they preferred each time, GPT-Live or Advanced Voice Mode.

Overall preferenceTurn-takingHandling interruptionsFluencyNaturalness

Head-to-head preference rate (50% = tied with AVM):

GPT-Live-1
75.7%
GPT-Live-1 mini
69.2%

There's also a separate conversation-rating test scored independently (out of 7 points, covering conversational fluency and overall comfort):

ModelConversational fluencyComfort
GPT-Live-14.965.19
GPT-Live-1 mini4.334.47
Advanced Voice Mode3.803.82

These three general-purpose benchmarks weren't built for voice specifically — they test the model's underlying capability, again compared against Advanced Voice Mode:

GPQA
Tests expert-level scientific reasoning across biology, chemistry, and physics.
BrowseComp
Tests the ability to find hard-to-locate information on the web (agentic web search).
τ³-Voice Telecom
An OpenAI-built internal benchmark that has voice agents handle real, multi-turn phone-support tasks.

GPQA accuracy, broken down by backend reasoning-effort tier:

AVM
45.3%
GPT-Live-1 mini
74.9%
GPT-Live-1 · Instant
76.5%
GPT-Live-1 · Medium
81.7%
GPT-Live-1 · High
84.2%

BrowseComp accuracy:

AVM
0.7%
GPT-Live-1 mini
31.6%
GPT-Live-1 · Instant
35.1%
GPT-Live-1 · Medium
60.6%
GPT-Live-1 · High
75.2%

τ³-Voice Telecom tracks both task-completion rate and time taken — higher reasoning effort means more accurate but slower:

ModelMedian task timeTask success rate
AVM385.5 sec29.5%
GPT-Live-1 mini290.9 sec39.5%
GPT-Live-1 · Instant231 sec37.3%
GPT-Live-1 · Medium295 sec59.6%
GPT-Live-1 · High386 sec63.4%
Which backend engine: GPT-Live-1 (Instant) and mini both use GPT-5.5 Instant on the backend; GPT-Live-1 Medium and High use GPT-5.5 Thinking, at medium and high reasoning-effort settings respectively. The τ³ benchmark was run using a custom user model driven by OpenAI's latest reasoning model. All figures are OpenAI's own reported data.
6In Practice

Open ChatGPT Voice — here's what's actually different now

More than 150 million people use ChatGPT's voice and dictation features every week — as a hands-free daily assistant, for practicing a language, for bedtime stories, or just to chat during a commute. Starting today, tapping the voice button means you're using GPT-Live.

Official OpenAI demo video (Chinese/English captions): several OpenAI staffers discuss the model, with two real-world scenes worked in — dictating hands-free edits to an offer-letter reply, and, mid-dance, using voice to look something up without touching a phone. Source: OpenAI, "Introducing GPT-Live."
Official OpenAI demo video (Chinese/English captions): the actual startup experience after tapping the ChatGPT voice button. Source: OpenAI, "Introducing GPT-Live."
150M+ / week
people who use ChatGPT voice and dictation weekly
2 versions
GPT-Live-1 (default for Go / Plus / Pro) and mini (default for Free)
9 voices
all remastered for this launch
GPT-5.5
the frontier model handling deep tasks in the background (Instant / Thinking)

More like talking to a real person, and better at listening too

You can jump in with a question at any time, pause to gather your thoughts, or ask it to slow down. It'll use responses like "mm-hm" or "got it" to show it's following along. All nine voices have been remastered as well.

Old version · When you pause

Even a brief pause to find the right words is likely to be read as "you're done," and it cuts in and interrupts you.

GPT-Live · When you pause

It waits quietly, without rushing to jump in. Tell it to stay quiet and just listen, and it will. In noisy settings — traffic, other people talking — it also stays better focused on your voice.

Smarter answers when you need them

Voice can now call on the latest frontier model, and you can choose the reasoning effort yourself: pick Instant for speed, or Medium/High when you want it to think harder.

Instant
Fastest replies — good for everyday quick questions.
Medium
Takes a bit more time to think, balancing speed and depth.
High
For when you want it to really think it through — longest thinking time.
Official OpenAI demo video: the actual interface for choosing reasoning effort (Instant / Medium / High) during a voice conversation. Source: OpenAI, "Introducing GPT-Live."

Some answers are more useful when you can see them

While it's talking, ChatGPT can now show you cards directly — for weather, stocks, sports scores, and more — without needing to type anything to confirm. Voice still supports search, memory, and image/file uploads as before.

Official OpenAI demo video (Chinese/English captions): while packing a suitcase, someone asks "can you show me this week's weather in Mexico City" — ChatGPT pops up a weather card directly, no need to stop and type to confirm. Source: OpenAI, "Introducing GPT-Live."

Two more official screenshots from OpenAI:

ChatGPT Voice showing a card with upcoming international soccer fixtures
Sports fixtures card example
ChatGPT Voice showing a map card of nearby bars
Map card example (nearby bars to watch the game)
7Safety

Safety design built specifically for voice

On top of the latest model's baseline safety, GPT-Live adds dedicated safety training and protections tailored to voice as a new medium.

Safety testing that's closer to real-world use

OpenAI extended its safety testing to audio-native evaluations, built a batch of tests using synthetic audio, and had internal experts run red-teaming focused specifically on voice-related risks. Per the original post, GPT-Live matched or beat Advanced Voice Mode across nearly every evaluation area. The key risk areas covered include:

Self-harmPsychosis and maniaEmotional dependence on AIViolenceSexual content

Protections that can step in mid-speech

Because voice conversation happens in real time, OpenAI built protections that can act while the model is still speaking. If the system detects potentially unsafe content, it can do one of three things:

Action 1
Steer the model toward a safer response
Action 2
Add extra safety guidance or support resources
Action 3
For higher-risk cases, end the voice conversation directly

For conversations involving self-harm, OpenAI adapted ChatGPT's support pathways for the voice version, including expert-reviewed crisis-hotline support.

Teen protections and voice safeguards

For teen users, the model's training includes age-appropriate behavior. Parents can use Parental Controls to decide whether their teen can use ChatGPT Voice at all; in high-risk situations showing signs of self-harm or suicidal intent, a linked parent may be notified. In addition, GPT-Live only uses a fixed set of preset voices, built with protections in place and not designed to mimic real people's voices.

8Rollout

Can you use it now? Availability and current limitations

GPT-Live is rolling out to ChatGPT users worldwide starting now, covering iOS, Android, and ChatGPT.com. Here's a breakdown of which version each plan defaults to.

GPT-Live-1Default for Go / Plus / Pro
The standard version, now the default voice model for Go, Plus, and Pro users.
GPT-Live-1 miniDefault for Free
The lightweight version, now the default voice model for Free users.
Legacy Standard / Advanced Voice Mode
Still available to switch to manually. Features GPT-Live doesn't yet support — like voice paired with video or screen sharing — can still be used there.

A few current limitations worth knowing: GPT-Live has been optimized for ChatGPT's most-used languages, and some languages may still come with a non-native accent or reduced fluency — OpenAI says it's working on improvements. At launch, it doesn't yet support voice combined with video or screen sharing, though those are coming soon. An API is also on the way; developers and businesses can sign up on a waitlist form to be notified.

  • For everyday voice Q&A, language practice, or chatting during a commute, interruptions, thinking pauses, and requests to slow down are now recognized naturally — no more getting talked over or misread as often.
  • For web lookups, expert-level science questions, or complex multi-step tasks, voice conversation can now reach accuracy close to a deep-reasoning model, since GPT-5.5 is handling it in real time behind the scenes.
  • For weather, stocks, sports scores, and similar lookups, voice conversation now shows visual cards directly — no need to type to confirm.
  • Parents can use Parental Controls to manage whether teens can use voice features, and high-risk conversations involving self-harm or suicidal intent will notify a linked parent.
GPT-Live is built on a full-duplex architecture, meaning it can listen and speak at the same time. For questions that need web search, deeper reasoning, or more complex work, it delegates to our latest frontier model behind the scenes, then brings the result back into the conversation once it's ready. OpenAI, "Introducing GPT-Live," July 8, 2026
This article is based on OpenAI's official launch blog post, "Introducing GPT-Live" (July 8, 2026). The benchmark data, preference results, and safety-testing conclusions cited here are OpenAI's own reported figures and have not been independently verified by a third party. Original post: openai.com/index/introducing-gpt-live/