Latency numbers mean nothing until you hear them. This simulator plays the same exchange at different total delays and breaks the budget down by hop, so you can tell where a pipeline is spending its time and at what point a caller starts talking over the agent.
Under about 700 milliseconds the conversation feels normal. Past a second, callers start repeating themselves and talking over the agent, which cascades because barge-in handling then has to be perfect. The failure is not that it is slow; it is that slowness breaks turn-taking entirely.
Almost always the reasoning hop, and almost always fixed by shortening the response rather than by changing model. Streaming synthesis from the first sentence buys the most perceived improvement for the least engineering. Swapping speech vendors is the usual first attempt and the least effective one.