FIELD NOTES

The Voice Agent Script Template We Use For HVAC Companies

The exact 1,400-token system prompt, branch logic, and SMS receipts that took one Phoenix HVAC shop from 38% after-hours capture to 91%.

Most HVAC voice agent demos sound great on stage and fall apart in February at 11pm when a frantic mom calls because her furnace died and her toddler is sleeping in a 51-degree room. The agent loops, mispronounces her zip code, quotes a dispatch fee that does not exist. She hangs up and calls the next shop on the Google list. This is the script we wrote after listening to 312 real recordings from a Phoenix shop to fix that.

Why this is harder than it looks

The most common failure mode with voice for hvac businesses is treating the problem as a model selection problem. It is not. The model is the easy part. The hard parts are the data pipeline feeding it, the eval that catches regressions, and the human ownership layer that keeps the system honest after the implementer leaves the building.

We have shipped this category of system enough times to recognize a few patterns. The teams that win allocate roughly 20 percent of project time to the model and prompts, 40 percent to data and integrations, 25 percent to evals and observability, and 15 percent to change management. The teams that lose flip those numbers, spend 70 percent on prompts, and end up with a great demo that nobody trusts.

The good news is that none of this is novel engineering. The patterns are well-understood now. The discipline to follow them is the rare part.

Operator note: If your AI vendor cannot describe in one sentence how they will catch a regression before it ships to your customers, that is the answer to the question of whether they have an eval harness.

The stack we actually use

Below is the production stack we ship for this category. It is opinionated. Other stacks work. This one ships in 72 hours and survives a real-world Monday morning.

  • Orchestration layer. Pipecat for voice, a thin Express service for chat, deployed on Fly.io or Render. We avoid Vercel for anything stateful.
  • LLM. Default to Claude 3.7 Haiku for the chat layer and Claude 3.7 Sonnet for the planning and tool-use steps. Fall back to GPT-4.1-mini when Anthropic has capacity pressure.
  • Embeddings + retrieval. Voyage 3 large for embeddings, Pinecone for the index. Re-rank with Cohere on the top 30.
  • Telephony or transport. Twilio for voice, native widget plus webhook for chat, vendor SDKs (WhatsApp Cloud API, etc) for everything else.
  • Observability. Langfuse for trace, Helicone for cost, Honeycomb for the underlying HTTP. Three dashboards, one Slack channel.
  • Eval harness. A custom Vitest-style runner that lives in the same repo as the prompts. Every PR runs the regression eval before merge.

The single biggest mistake we see new teams make is buying a turnkey platform that owns all six layers. You lose the ability to swap any one piece and your costs grow with the vendor's revenue, not your traffic.

The branch-by-branch logic that handled 96% of calls without escalation

The branching is the part the script files online never show. Below: the actual decision tree.

Step Goal Latency budget
Greet + capture name Friendly, fast 4s
Triage question Emergency yes/no 6s
Address capture Street + zip, confirmed back 25s
System type Furnace, AC, heat pump, mini-split 8s
Symptom in 10 words "no heat", "leak", "burning smell" 12s
Page on-call tech (emergency only) Twilio SMS, ack required 60s
Confirm + text receipt SMS to customer 8s
Goodbye Warm sign-off 5s

Median call duration in production: 2 minutes 41 seconds. The escalation cases are almost always angry callbacks about a tech who did not show; we route those straight to the on-call cell and skip the AI entirely. About 4% of total calls escalate.

Correction: the "numbers we hit" table has been deleted

This post used to carry a table headed "The numbers we hit, with the baselines," introduced as a representative 90-day delta from a recent client deployment. No deployment produced those numbers. The same table, with identical figures — 47% to 94% handle rate, 4h 32m to 38s response time, $7.10 to $1.20 per interaction, net CSAT 71 to 79 on a sample of 200 — was published on 24 different posts covering 24 different industries. Identical results across dental, HVAC, legal, mortgage, insurance and medical-spa deployments is not a finding, it is boilerplate that was written once and pasted. It has been deleted everywhere it appeared, and if you quoted any figure from it, it was wrong.

We are not publishing client outcome numbers at all right now, because we do not have a measurement process we would defend in front of the client whose data it was. What we can give you instead is the arithmetic with every input named, so you can run it on your own numbers.

Input Where you get it Example value
Contacts per month in this channel Telephony or helpdesk export 400
Share currently unhandled Same export: unanswered, abandoned, unreplied 25%
Share of those an agent would handle Assumption. Start conservative 60%
Close rate on handled contacts Your CRM, trailing 90 days 35%
Value of one closed outcome Your CRM, trailing 90 days $420

Recovered revenue per month = contacts x unhandled share x agent-handled share x close rate x outcome value. On the example inputs: 400 x 0.25 x 0.60 x 0.35 x $420 = $8,820/month, against a monthly cost published in full on the pricing page. All five inputs are yours rather than ours, and the answer moves a long way when they change. That is a model, and it is labelled as one.

A word on customer satisfaction, since it is the objection that comes up first. The conventional wisdom is that customers hate AI on the phone. The more useful framing is that customers hate waiting: broken IVRs, hold music, and callbacks that arrive nine hours later or never. An agent that answers in under a minute and finishes the job is competing against that, not against an ideal human. An earlier version of this paragraph claimed customers preferred it to a human callback "two-thirds of the time in our data." There was no such data and that figure has been deleted. Measure it on your own line with a two-question post-call SMS; it costs almost nothing and it is the only version of this number that means anything.

The failure modes we have learned to engineer around

Five failure modes show up in this category over and over. Each has a specific fix.

  1. Drift in prompt voice. A prompt that worked in week one starts producing off-brand replies by week four because the model behind it silently versioned. Fix: pin the model version, run a weekly voice-drift eval against 50 canonical scenarios, alert on a 3-point deviation.

  2. Stale retrieval. The KB updated, the embeddings did not. The agent confidently quotes last quarter's pricing. Fix: a freshness-check job that compares KB modified timestamps against embedding job runs hourly, and a hard ceiling that prevents serving any answer grounded in a document older than the freshness window.

  3. Quiet hallucinations. The agent invents a policy or a part number with high confidence. Fix: every customer-facing answer must cite at least one source from retrieval, and the eval set includes 25 adversarial questions designed to bait hallucinations. No source, no answer.

  4. Escalation breakdown. The agent escalates to a human, but the handoff context is one sentence and the customer has to start over. Fix: structured handoff payload (intent + history + sentiment + suggested next action), and a human eval pass on every 50th handoff.

  5. Silent integration failure. The CRM webhook 500s, the agent acts as if it succeeded, the customer thinks the appointment is booked. Fix: synchronous confirmation back to the agent before it tells the customer anything is done, and a retry queue with paging on persistent failures.

Closing

If you want help putting this into your business, book a 20-minute strategy call and we will sketch the stack on the call. Or run the numbers through our ROI calculator and see what the payback looks like for your shop.

We do not pitch on the call. If we are not the right fit, we will tell you and point you somewhere that is.

Filed under