What shipped

OpenAI released GPT-Live-1 in its API on 10 September, describing it as bringing “natural, full-duplex voice conversations to the API, with stronger instruction following, custom voices, and telephony support”, according to the company’s developer announcement. It is priced at $0.05 per minute.

Full duplex is the whole claim. The model can listen while it is speaking, and react mid-sentence to interruptions, acknowledgements, pauses, changes of direction and background speech while its own audio is still flowing. The previous generation, gpt-realtime-2.1, took turns: it finished, then listened.

Why turn-taking was the bottleneck

Anyone who has used a voice assistant knows the failure mode. You start to correct it, it keeps talking over you, and you wait for a gap. Half-duplex models handle that by detecting silence, which is why they misread a thoughtful pause as the end of your sentence and a filler word as a new instruction.

A close-up of a studio microphone
Full-duplex models can listen while speaking, rather than waiting for silence. RDNE Stock project · pexels · Pexels License

Removing the constraint changes what a voice agent can be used for. A support call where the caller says “no, wait, the other order” three words into an answer is normal human behaviour and, until now, an awkward case to engineer around.

The numbers, and who reported them

OpenAI says the model shows roughly a 30-percentage-point improvement over gpt-realtime-2.1 on Full Duplex Bench, and that paired with GPT-6 Astra at medium reasoning effort it places first on the Tau3 evaluations, as summarised by Data Studios. Both figures are OpenAI’s own, on its own model, and there is no independent reproduction yet.

What it costs in practice

The $0.05 per minute covers the voice layer and is billed per second rather than rounded up. Backend calls to the Responses API and any tool use are charged at their normal rates, so a voice agent that reasons with a frontier model and calls three tools per turn does not cost five cents a minute in total.

A developer working at a computer
Backend reasoning calls and tool use are billed separately from the voice layer. cottonbro studio · pexels · Pexels License

The release also includes telephony support and an expanded set of voices covering more accents, dialects and languages. Custom voices require a separate eligibility process.

What to watch

The claim that needs independent testing is latency under load: full-duplex behaviour that holds in a demo can degrade when the model is also calling tools and waiting on a backend. The second thing is what developers do with delegation — GPT-Live-1 is designed to hand reasoning to a separate backend model, which means the interesting question is not how well it talks but how well it waits.