FIELD NOTES

The Voice Agent Tech Stack: Twilio, Deepgram, Cartesia, Anthropic

Our exact production voice stack. Every component, every version, every config knob, every fallback. Steal it.

Below is the exact production voice stack we ship to clients. Every component, every version, every config knob, every fallback path. It is opinionated. It works.

Why this is harder than it looks

The most common failure mode with voice for stack businesses is treating the problem as a model selection problem. It is not. The model is the easy part. The hard parts are the data pipeline feeding it, the eval that catches regressions, and the human ownership layer that keeps the system honest after the implementer leaves the building.

We have shipped this category of system enough times to recognize a few patterns. The teams that win allocate roughly 20 percent of project time to the model and prompts, 40 percent to data and integrations, 25 percent to evals and observability, and 15 percent to change management. The teams that lose flip those numbers, spend 70 percent on prompts, and end up with a great demo that nobody trusts.

The good news is that none of this is novel engineering. The patterns are well-understood now. The discipline to follow them is the rare part.

Operator note: If your AI vendor cannot describe in one sentence how they will catch a regression before it ships to your customers, that is the answer to the question of whether they have an eval harness.

The stack we actually use

Below is the production stack we ship for this category. It is opinionated. Other stacks work. This one ships in 72 hours and survives a real-world Monday morning.

  • Orchestration layer. Pipecat for voice, a thin Express service for chat, deployed on Fly.io or Render. We avoid Vercel for anything stateful.
  • LLM. Default to Claude 3.7 Haiku for the chat layer and Claude 3.7 Sonnet for the planning and tool-use steps. Fall back to GPT-4.1-mini when Anthropic has capacity pressure.
  • Embeddings + retrieval. Voyage 3 large for embeddings, Pinecone for the index. Re-rank with Cohere on the top 30.
  • Telephony or transport. Twilio for voice, native widget plus webhook for chat, vendor SDKs (WhatsApp Cloud API, etc) for everything else.
  • Observability. Langfuse for trace, Helicone for cost, Honeycomb for the underlying HTTP. Three dashboards, one Slack channel.
  • Eval harness. A custom Vitest-style runner that lives in the same repo as the prompts. Every PR runs the regression eval before merge.

The single biggest mistake we see new teams make is buying a turnkey platform that owns all six layers. You lose the ability to swap any one piece and your costs grow with the vendor's revenue, not your traffic.

The exact stack, version-pinned

  • Telephony: Twilio Programmable Voice. Number forwarded after hours. Concurrency ceiling raised to 50 per client.
  • Media handoff: Twilio Media Streams over WebSocket, mu-law 8kHz.
  • Orchestrator: Pipecat 0.0.49. Deployed on Fly.io 1x cpu, 1GB ram, region matching telephony POP.
  • STT: Deepgram Nova-3 streaming. endpointing_ms tuned per vertical (HVAC: 800, dental: 600).
  • VAD: WebRTC VAD on the inbound side, agent-side VAD off (Pipecat handles it).
  • LLM: Anthropic Claude 3.7 Haiku for main loop, Claude 3.7 Sonnet for tool selection. Streaming on.
  • TTS: Cartesia Sonic. Custom voice clone per shop (with permission). Output at 24kHz, resampled to 8kHz for the call.
  • Tool calls: ServiceTitan or Housecall Pro for HVAC; Mindbody for med spa; Clio for legal. Each tool wrapped in a retry-with-backoff and an idempotency key.
  • Observability: Langfuse + Helicone + Honeycomb.
  • Recording + transcript: S3 in client account, 90-day retention, redacted transcript in our DB.

Fallback paths: if Anthropic is unavailable, route to OpenAI GPT-4.1-mini. If both are unavailable, fall back to a static IVR with voicemail-to-text.

Total cost per minute on this stack: $0.13 to $0.18 depending on TTS engagement and tool-call frequency.

The failure modes we have learned to engineer around

Five failure modes show up in this category over and over. Each has a specific fix.

  1. Drift in prompt voice. A prompt that worked in week one starts producing off-brand replies by week four because the model behind it silently versioned. Fix: pin the model version, run a weekly voice-drift eval against 50 canonical scenarios, alert on a 3-point deviation.

  2. Stale retrieval. The KB updated, the embeddings did not. The agent confidently quotes last quarter's pricing. Fix: a freshness-check job that compares KB modified timestamps against embedding job runs hourly, and a hard ceiling that prevents serving any answer grounded in a document older than the freshness window.

  3. Quiet hallucinations. The agent invents a policy or a part number with high confidence. Fix: every customer-facing answer must cite at least one source from retrieval, and the eval set includes 25 adversarial questions designed to bait hallucinations. No source, no answer.

  4. Escalation breakdown. The agent escalates to a human, but the handoff context is one sentence and the customer has to start over. Fix: structured handoff payload (intent + history + sentiment + suggested next action), and a human eval pass on every 50th handoff.

  5. Silent integration failure. The CRM webhook 500s, the agent acts as if it succeeded, the customer thinks the appointment is booked. Fix: synchronous confirmation back to the agent before it tells the customer anything is done, and a retry queue with paging on persistent failures.

The 72-hour deploy plan

The build sprint below runs on a 72-hour clock. That is the engineering, not the engagement: our published promise is live in 21 days from kickoff, which wraps this sprint in scoping, evals, shadow mode and cutover. If you see "72 hours" and "21 days" on this site and wonder which is true, both are — one is the part where code gets written.

Hour 0-8. Kickoff. Interview the two people who do this job today. Pull 50 sample inputs (calls, chats, tickets). Establish baseline metrics. Identify the three top customer intents.

Hour 8-24. First-pass prompt. Wire the orchestration. Stand up the eval harness with 25 cases drawn from the sample inputs. The eval harness has to exist before the first prompt does.

Hour 24-40. Integrations. CRM webhook, calendar booking, payment link if relevant. Each integration ships with a synchronous confirmation path.

Hour 40-56. Internal QA. The two people we interviewed in hour 0 spend 90 minutes running the agent through their hardest scenarios. Their feedback drives the second-pass prompt.

Hour 56-68. Shadow traffic. Real customer interactions, AI answers, human reviews before the answer is sent. We are looking for any case where the AI's draft is worse than the human's draft.

Hour 68-72. Cutover. We flip the routing rule, monitor for the first hour, and hand off the on-call rotation to the client's champion. The implementer stays on standby for 7 days.

Operator note: The 72-hour clock is real but it assumes the client has decided on success criteria before we start. If success criteria are unclear at hour 0, the clock does not start until they are. This is the single biggest cause of pilot drift we see.

Next step

If you want to pressure-test the numbers above against your business, book a call. It is free and we do not bring slides.

Prefer to read more first? Our case studies walk through three full deployments: what worked, what we would do differently, and what each one cost.

Filed under