What shipped
OpenAI released GPT-Live-1 in its API on 10 September, describing it as bringing “natural, full-duplex voice conversations to the API, with stronger instruction following, custom voices, and telephony support”, according to the company’s developer announcement. It is priced at $0.05 per minute.
Full duplex is the whole claim. The model can listen while it is speaking, and react mid-sentence to interruptions, acknowledgements, pauses, changes of direction and background speech while its own audio is still flowing. The previous generation, gpt-realtime-2.1, took turns: it finished, then listened.
Why turn-taking was the bottleneck
Anyone who has used a voice assistant knows the failure mode. You start to correct it, it keeps talking over you, and you wait for a gap. Half-duplex models handle that by detecting silence, which is why they misread a thoughtful pause as the end of your sentence and a filler word as a new instruction.

Removing the constraint changes what a voice agent can be used for. A support call where the caller says “no, wait, the other order” three words into an answer is normal human behaviour and, until now, an awkward case to engineer around.
The numbers, and who reported them
OpenAI says the model shows roughly a 30-percentage-point improvement over gpt-realtime-2.1 on Full Duplex Bench, and that paired with GPT-6 Astra at medium reasoning effort it places first on the Tau3 evaluations, as summarised by Data Studios. Both figures are OpenAI’s own, on its own model, and there is no independent reproduction yet.
What it costs in practice
The $0.05 per minute covers the voice layer and is billed per second rather than rounded up. Backend calls to the Responses API and any tool use are charged at their normal rates, so a voice agent that reasons with a frontier model and calls three tools per turn does not cost five cents a minute in total.

The release also includes telephony support and an expanded set of voices covering more accents, dialects and languages. Custom voices require a separate eligibility process.
What to watch
The claim that needs independent testing is latency under load: full-duplex behaviour that holds in a demo can degrade when the model is also calling tools and waiting on a backend. The second thing is what developers do with delegation — GPT-Live-1 is designed to hand reasoning to a separate backend model, which means the interesting question is not how well it talks but how well it waits.