FIELD NOTES

Voice Agent Latency: A Deep Dive Into What Actually Matters

Time-to-first-audio is the only latency number that moves bookings. Here is how to measure it, what is broken about p95, and the 7 places we shave milliseconds.

How the post opens

Latency is the single most-cheated number in voice agent demos. Vendors quote p95 to hide their tail. They benchmark on the demo POP and run you on the cheap one. They start the clock from the wrong place. Below is the only latency number that moves bookings and how to actually measure it.

What it argues

The numbers we hit, with the baselines Below is a representative 90-day delta from a recent client deployment in this category. The baseline is from the trailing 90 days before deploy, the post numbers are from the 90 days after first traffic. Metric Baseline After 90 days Delta --- --- --- --- Inbound handle rate 47% 94% +47pts Median time to first response 4h 32m 38s -99% Conversion to booked outcome 18% 31% +13pts Cost per resolved interaction $7.10 $1.20 -83% 5-star reviews per month 6 14 +133% Net CSAT (sample of 200) 71 79 +8pts Numbers vary by vertical. Once you get the data pipeline right and the eval gate in place, handle rate jumps 40-60 percentage points within the first month, conversion follows two to four weeks later as the prompts get tuned, and cost-per-interaction drops as caching and model routing kick in. The CSAT delta is the one that surprises new operators.

Filed under

Why this one exists

Time-to-first-audio is the only latency number that moves bookings. Here is how to measure it, what is broken about p95, and the 7 places we shave milliseconds. It is filed under AI Voice because that is where operators looking for this problem actually start, and it is written from production work rather than from a content calendar.