FIELD NOTES

Voice Agent Latency: A Deep Dive Into What Actually Matters

Time-to-first-audio is the only latency number that moves bookings. Here is how to measure it, what is broken about p95, and the 7 places we shave milliseconds.

Latency is the single most-cheated number in voice agent demos. Vendors quote p95 to hide their tail. They benchmark on the demo POP and run you on the cheap one. They start the clock from the wrong place. Below is the only latency number that moves bookings and how to actually measure it.

Correction: the "numbers we hit" table has been deleted

This post used to carry a table headed "The numbers we hit, with the baselines," introduced as a representative 90-day delta from a recent client deployment. No deployment produced those numbers. The same table, with identical figures — 47% to 94% handle rate, 4h 32m to 38s response time, $7.10 to $1.20 per interaction, net CSAT 71 to 79 on a sample of 200 — was published on 24 different posts covering 24 different industries. Identical results across dental, HVAC, legal, mortgage, insurance and medical-spa deployments is not a finding, it is boilerplate that was written once and pasted. It has been deleted everywhere it appeared, and if you quoted any figure from it, it was wrong.

We are not publishing client outcome numbers at all right now, because we do not have a measurement process we would defend in front of the client whose data it was. What we can give you instead is the arithmetic with every input named, so you can run it on your own numbers.

Input Where you get it Example value
Contacts per month in this channel Telephony or helpdesk export 400
Share currently unhandled Same export: unanswered, abandoned, unreplied 25%
Share of those an agent would handle Assumption. Start conservative 60%
Close rate on handled contacts Your CRM, trailing 90 days 35%
Value of one closed outcome Your CRM, trailing 90 days $420

Recovered revenue per month = contacts x unhandled share x agent-handled share x close rate x outcome value. On the example inputs: 400 x 0.25 x 0.60 x 0.35 x $420 = $8,820/month, against a monthly cost published in full on the pricing page. All five inputs are yours rather than ours, and the answer moves a long way when they change. That is a model, and it is labelled as one.

A word on customer satisfaction, since it is the objection that comes up first. The conventional wisdom is that customers hate AI on the phone. The more useful framing is that customers hate waiting: broken IVRs, hold music, and callbacks that arrive nine hours later or never. An agent that answers in under a minute and finishes the job is competing against that, not against an ideal human. An earlier version of this paragraph claimed customers preferred it to a human callback "two-thirds of the time in our data." There was no such data and that figure has been deleted. Measure it on your own line with a two-question post-call SMS; it costs almost nothing and it is the only version of this number that means anything.

The failure modes we have learned to engineer around

Five failure modes show up in this category over and over. Each has a specific fix.

  1. Drift in prompt voice. A prompt that worked in week one starts producing off-brand replies by week four because the model behind it silently versioned. Fix: pin the model version, run a weekly voice-drift eval against 50 canonical scenarios, alert on a 3-point deviation.

  2. Stale retrieval. The KB updated, the embeddings did not. The agent confidently quotes last quarter's pricing. Fix: a freshness-check job that compares KB modified timestamps against embedding job runs hourly, and a hard ceiling that prevents serving any answer grounded in a document older than the freshness window.

  3. Quiet hallucinations. The agent invents a policy or a part number with high confidence. Fix: every customer-facing answer must cite at least one source from retrieval, and the eval set includes 25 adversarial questions designed to bait hallucinations. No source, no answer.

  4. Escalation breakdown. The agent escalates to a human, but the handoff context is one sentence and the customer has to start over. Fix: structured handoff payload (intent + history + sentiment + suggested next action), and a human eval pass on every 50th handoff.

  5. Silent integration failure. The CRM webhook 500s, the agent acts as if it succeeded, the customer thinks the appointment is booked. Fix: synchronous confirmation back to the agent before it tells the customer anything is done, and a retry queue with paging on persistent failures.

The 7 places we shave milliseconds

Latency optimization is mostly boring infrastructure. Here are the seven places we have shaved real time, in priority order.

  1. Endpointing tuned to the vertical. Default VAD waits 600ms for silence. HVAC callers pause to think; legal callers pause less. Tune per vertical.
  2. Streaming STT not batch. Deepgram Nova-3 streaming gets you transcript chunks at 100-200ms intervals; batch waits for utterance complete.
  3. First-token streaming on the LLM. Anthropic's streaming API lets you start TTS on the first sentence boundary, not the full response.
  4. TTS engine choice. Cartesia Sonic is ~120ms TTFA; ElevenLabs Turbo is ~280ms; the difference is audible.
  5. POP geography. Telephony POP within 50ms of your LLM POP. Twilio Frankfurt + Anthropic Frankfurt = 18ms RTT. Mix coasts and you eat 80ms per hop.
  6. Co-located orchestration. Pipecat in the same region as both above, not on a $5/mo VPS in Singapore.
  7. Skip the planner on common intents. 60% of inbound calls have one of three intents. Match-and-respond without a planning LLM call drops 400ms.

Combined, these seven took our p50 TTFA from 1,420ms to 540ms. p99 from 4,100ms to 1,200ms. The booking rate ticked up 9% on the same prompts.

What to measure in the first 30 days

Most teams measure too many things and then measure nothing. The 30-day measurement plan is short:

  • Handle rate. Of inbound contacts in the channel where AI is now answering, what percent did AI successfully complete versus escalate or drop. This is the deflection metric in chat language.
  • Time-to-outcome. Median minutes from first contact to whatever the business cares about: booked, ordered, refunded, qualified.
  • Cost per completed interaction. All-in, including telephony, STT, TTS, LLM, observability, and your eval and ops time amortized.
  • Brand-voice score. A weekly sample of 25 interactions, scored 1-5 by your marketing lead. Track the median and the bottom-quartile floor. The floor matters more than the median.
  • Escalation reason mix. Why are escalations happening, in 6-8 buckets, week over week. Anomalies here are leading indicators of prompt or KB issues.

Five metrics, one dashboard, one Monday review. If the metric is not on the dashboard, it does not exist for the first 30 days.

After 30 days you can add CSAT, conversion-to-revenue, and channel attribution. Adding them earlier just adds noise.

The tradeoffs we made and why

Every architecture choice in this category is a tradeoff. Here are the ones we have made consciously, and the alternative we did not pick.

We use Pipecat instead of building our own orchestration. The win is months of saved engineering. The cost is being a release behind on a few model integrations.

We use Anthropic as the default LLM with OpenAI failover, not the other way around. The win is consistently lower hallucination rates in our eval set. The cost is slightly higher cost-per-token at the equivalent model tier.

We run a custom eval harness instead of using a vendor product. The win is the eval set lives in the same Git repo as the prompts, so PRs that change prompts fail the build if they break an eval. The cost is we own the upkeep.

We use Twilio over a cheaper telephony provider. The win is concurrency ceiling and the depth of the diagnostic tools. The cost is roughly 18 percent more per minute.

We deploy on Fly.io, not Vercel. The win is real persistent connections and global edge POPs that survive long-lived voice sessions. The cost is more devops overhead.

These tradeoffs are not laws. They are starting points. If you tell us you have an existing GCP estate, we will adapt the stack. The principles do not move; the implementations do.

What now

If something here tripped a wire, grab time on the calendar. We do a free 20-minute working session where we map your current flow and circle the two places AI moves the needle this quarter.

You can also browse the rest of the blog for 40-plus posts in the same operator voice, no fluff.

Filed under