Stop trusting the headline minute rate. Here's the actual TTS / STT / LLM / telephony / orchestration math for a 2026 voice agent, what self-hosting really saves, and when fixed-fee beats usage-based pricing.
Pricing note, 5 September 2026. Our published prices changed on this date: Operators moved from $4,950 build / $1,997 per month to $7,500 / $2,950, and Agent in a Day moved from $497 one-time to $2,500 build / $497 per month. The arithmetic below was run against the prices in force when it was written and has been left as it was rather than quietly restated. Current prices are on the pricing page.
Every voice-AI vendor in 2026 quotes a single number: $0.12/min, $0.30/min, $0.99/min. The number is clean. The math behind it is not.
A voice agent on a phone call is five paid services running concurrently with margin layered on top. If you don't see the line items, you can't tell whether the platform fee is fair, whether self-hosting will actually save you anything, and whether your pricing to the customer makes sense.
I run TrainYourAgent. We have ~$20K/mo in recurring voice-agent revenue across HVAC, healthcare, legal, real estate, and ecommerce. This is the per-minute math we use to scope every deal — current as of Q2 2026, real providers, real rates.
A sixth line nobody quotes: failed calls + silence. You're paying for the carrier minute whether or not the agent says anything useful.
Baseline assumption: a 60-second billable minute. The caller talks for ~45 seconds, the agent talks for ~30 seconds, four turns, ~2,400 input tokens and ~1,200 output tokens to the LLM, ~600 characters of TTS.
| Component | Provider | May 2026 rate | Per-min cost |
|---|---|---|---|
| Telephony (inbound US local) | Twilio | $0.0085/min | $0.009 |
| STT (streaming) | Deepgram Nova-3 | $0.0043/min | $0.004 |
| LLM | GPT-4.1-mini ($0.40 / $1.60 per M tokens) | usage | $0.003 |
| TTS | ElevenLabs Flash v2.5 | $0.10 per 1k chars | $0.060 |
| Orchestration | Self-hosted Pipecat on $40/mo Fly machine | amortized | $0.008 |
| Total floor | ~$0.084 / min |
About 8.4 cents a minute if you stitch it together yourself with mid-tier components.
Now the same minute on a managed platform:
The gap between DIY ($0.08) and managed ($0.15–0.25) is your convenience tax. Under ~50,000 minutes/month, the convenience is worth it. Over that, you're leaving real money on the table.
People obsess over which LLM to use. The LLM is the cheapest line item. $0.003/min is a rounding error. Two costs actually matter:
TTS dominates. ElevenLabs Flash is 70% of variable cost in the breakdown above. Switching to Cartesia Sonic ($0.04/min equivalent) or Rime cuts it in half but you trade voice quality. For luxury hotels or high-end legal — pay the ElevenLabs premium. For HVAC after-hours triage — Cartesia is fine.
Telephony compounds at scale. $0.009/min sounds tiny until you're doing a million minutes/month. That's $9k just to keep the lines open. Negotiate volume rates with Twilio at 250k+ min — typically 20–35% off rack. Telnyx is ~25% cheaper out of the box but has more carrier-routing edge cases.
Three costs that always show up in your bill even when nobody mentions them on the sales call:
Add 12–15% on top of your happy-path cost to model real bills.
Want the spreadsheet I use to scope every deal? Grab the AI Voice Buyer's Guide — includes the per-minute calculator with all 2026 rates pre-filled.
The cutover point in our practice:
Self-hosting isn't free. Budget one senior engineer-week for the initial build, plus 4–6 hours/month maintenance. That's ~$3,000/mo loaded cost. Self-hosting breaks even against Vapi at roughly 30,000 min/mo.
Once you know your cost floor, the pricing model question matters. The math we use:
Fixed-fee (our own lanes are $99/mo self-serve, $2,950/mo Operators, $4,997/mo Scale) — best when:
Usage-based ($0.20/min or $X per booking) — best when:
Most SMB voice-agent shops should default to fixed-fee with a generous cap. That is what we do: Operators is $1,997/mo with 5,000 minutes included and $0.40/min after, on top of a $4,950 build fee. Published in full on the pricing page. A cap keeps the sales conversation about outcomes instead of unit economics, and it stops one high-volume customer eating the tier's margin.
The right question isn't "what does an AI voice agent cost in 2026." The right question is: what's the gross margin between your per-minute cost and what your customer charges per booking, dispatch, or recovered no-show?
A 60-second call that costs you 8 cents and books a $4,200 HVAC repair has 99.998% gross margin. That math doesn't change whether you're paying $0.08/min or $0.25/min on the input side. Pick the stack you can actually ship, ship it, and optimize the unit economics later — when you actually have units.
If you want me to scope this for your specific business, start a 7-day live trial or check the HVAC, healthcare, or ecommerce vertical pages.