This is the document an engineer would write after a diligence call, not a marketing diagram. Latency budget by hop, what happens when a provider degrades, how evals gate a deploy, what is logged and for how long, and the escape hatches that hand a live call to a human without dropping it.
The latency budget, hop by hop
The number that matters is time from caller silence to first audio out. Streaming speech recognition returns partials in under 300ms, the reasoning call is the variable term, and synthesis begins on the first sentence rather than the full response. Budget overruns are handled by shortening the response, not by making the caller wait.
Evals gate deploys, not the other way round
A scenario suite runs before any prompt or knowledge-base change reaches production: escalation cases, booking-conflict cases, refusal cases, and the ambiguity cases where a caller changes their mind. A regression blocks the deploy. Without this gate, prompt edits are guesses that only fail in front of customers.
Observability and retention
Every call produces a transcript, a structured summary, a disposition tag, and a latency trace. Traces are retained for debugging; transcripts follow the customer's retention policy, with redaction applied before logging under healthcare builds. You can pull any call and see exactly which model answered and how long each hop took.
Escape hatches
Warm transfer to a human, SMS handoff with the caller's details, and a queued callback are all first-class paths rather than error states. Low model confidence, an explicit request for a person, and any mention of legal, refund, or complaint all trigger one. The default is never to let an unhandled call die on the line.