A voice model built for agents, not demos
Amazon made Nova 2.5 Sonic generally available on 5 October, its latest speech-to-speech model for real-time voice agents. It is available through Amazon Bedrock in US East (N. Virginia), US West (Oregon), Europe (Stockholm) and Asia Pacific (Tokyo), at the same pricing as Nova 2 Sonic.
The feature list reads like a list of complaints about voice agents. Amazon cites improved reasoning, instruction following and tool-calling accuracy, lower latency, controllable turn-taking, and support for voice and text in the same session. The context window is 256,000 tokens, and the model speaks in seven languages with what Amazon calls expressive voices.

The detail that matters is asynchronous tool calling
Of everything in the release, asynchronous tool calling is the one that changes what a voice agent can do. A synchronous voice agent that needs to look something up goes quiet while it does: the caller hears silence, or a filler phrase, and the illusion of a conversation breaks. Calling the tool asynchronously lets the model keep speaking while the lookup runs.
The 256K context serves the same end. A voice agent handling a long support call needs to keep the whole call in view — what the caller said eight minutes ago, which account was verified, what was already tried — rather than re-asking. These are the two failure modes that have kept voice agents in demos and out of production queues.
What Amazon has not published is the latency figure. The announcement says lower latency without saying lower than what, by how much, or measured how. For a real-time voice product that is the number that decides whether it is usable, and its absence is conspicuous.

Priced to be adopted
Keeping the price at Nova 2 Sonic levels is the commercial signal. Amazon is not positioning this as a premium tier; it is making the newer model the default choice for anyone already building on the old one, which is how a cloud provider converts an installed base rather than selling to a new one.
The four launch regions are worth reading too. Stockholm and Tokyo alongside two US regions means Amazon is serving European and Japanese voice workloads from inside those jurisdictions — which matters for a product that processes recorded speech, the most personal data an enterprise agent routinely touches.
Amazon lists seven supported languages without naming them in the release. That, and the missing latency number, are the two things worth asking for before committing a call centre to it.