What does an AI receptionist actually cost per minute and per month?
A voice agent's cost is the sum of five components: telephony, speech-to-text, text-to-speech, the language model, and orchestration. Observed 2026 bands put a self-assembled stack at $0.05 to $0.15 per minute all-in, and a premium managed service as high as $0.40, which works out at roughly $0.30 to $0.50 per completed two-to-three-minute call.
What goes in, and what comes out
In: Calls per month and average call length
In: Per-minute rate for each of the five stack components
In: Any fixed platform or managed-service fee
In: The hourly cost of the staff time this would replace or relieve
Out: All-in cost per minute, built up component by component
Out: Cost per completed call
Out: Total monthly cost, usage plus fixed fee
Out: Side-by-side against the staffing cost you entered and against published self-serve plans
The method, stated
This is a bottom-up cost build. It sums five per-minute component rates, multiplies by the minutes you expect to consume, adds any fixed fee, and compares the total against a staffing cost you supply and against published self-serve list prices. It models the inputs you entered, not a quote.
monthly_minutes = calls_per_month x average_call_minutes
usage_cost = per_minute_all_in x monthly_minutes
monthly_total = usage_cost + fixed_monthly_fee
cost_per_call = per_minute_all_in x average_call_minutes + (fixed_monthly_fee / calls_per_month)
human_monthly = staff_hourly_cost x staffed_hours_per_week x 4.33
Every assumption this model bakes in
Component sliders default to the midpoint of the observed 2026 bands published on this page, which totals $0.13 a minute. Every one is editable, and the band is shown next to each slider so you can see where you have placed yourself. Orchestration is open-ended at the top, so its slider runs to $0.15 while its default sits at $0.055, the midpoint of the two stated bounds.
Billing is assumed to be per-minute with per-second or sub-minute granularity. Providers that round every call up to a full minute will cost more than this model shows on short calls.
Failed, abandoned and sub-ten-second calls still consume telephony and speech-to-text. The model does not subtract them, so it is slightly conservative on a line with heavy hang-ups.
Language-model cost is expressed per conversation-minute rather than per token, because per-minute is the unit the rest of the stack bills in. Token-heavy configurations, long system prompts and large retrieved context all push this line toward the top of its band.
Orchestration covers the platform layer, telephony routing, call state, integrations, logging and the vendor's margin. It is the widest band in the stack and the one that most separates a self-assembled build from a managed service.
The staffing comparison uses the fully loaded hourly cost you type in. We publish no salary figure, because a receptionist's loaded cost varies by market by a factor of several and any number we printed would be wrong for most readers.
What this model cannot tell you
It prices minutes, not outcomes. A cheaper agent that fails to book the appointment is not cheaper.
It does not include one-off build, integration, prompt and evaluation work, which for a production deployment is usually larger than the first year of usage cost.
It cannot tell you your real average call length. Pull that from your phone system; it is the single input the total is most sensitive to.
Self-serve list prices are entry-tier list prices and usually cap included minutes. The comparison is a starting point, not a like-for-like.
Per-minute component costs for a voice agent, observed 2026 bands
These are observed 2026 market bands, published so you can locate yourself inside them. They are not TrainYourAgent quotes, and nothing here is a promise about what your project will cost.
Telephony: $0.0085 – $0.014 — Inbound number, carrier minutes, call routing and media transport.
Speech-to-text (STT): $0.0025 – $0.016 — Streaming transcription of the caller. The spread reflects model quality and whether you self-host.
Text-to-speech (TTS): $0.015 – $0.030 — The agent's voice. Premium expressive voices sit at the top of the band.
Language model (LLM): $0.010 – $0.060 — Reasoning per conversation-minute. Long prompts and retrieved context push this up.
Orchestration and margin: $0.010 – $0.100+ — Platform layer, call state, integrations, logging, support and vendor margin. Widest band in the stack.
All-in per-minute and per-call cost, observed 2026 bands
These are observed 2026 market bands, published so you can locate yourself inside them. They are not TrainYourAgent quotes, and nothing here is a promise about what your project will cost.
Self-assembled stack: $0.05 – $0.15 per minute — You wire telephony, STT, TTS and an LLM together and operate it yourself.
Premium managed service: up to $0.40 per minute — Someone else owns the stack, the integrations, the evaluation loop and the pager.
Per completed call: $0.30 – $0.50 — A typical two-to-three-minute booking or triage call, priced end to end.
Published entry-tier monthly prices, self-serve AI receptionist market
Publicly listed entry-tier prices as observed in 2026. Included minutes, overage rates and feature caps differ substantially between them, so these are not like-for-like.
Rosie: $49 / mo — Entry tier list price.
Goodcall: $59 / mo — Entry tier list price.
myaifrontdesk: $65 / mo — Entry tier list price.
Smith.ai: $95 / mo — Entry tier list price.
What are the five components of AI receptionist cost?
Every voice agent, whoever sells it to you, is assembled from the same five parts, and every price you are ever quoted is those five parts plus a margin. Knowing the parts is what turns a quote into something you can evaluate. Telephony carries the call. Speech-to-text turns the caller's audio into words. A language model decides what to say. Text-to-speech turns that back into audio. Orchestration is everything that holds the loop together: call state, barge-in handling, function calls out to your calendar or CRM, logging, retries, and the vendor's margin. The observed 2026 bands for each are published in the table above. Add the floors and you land at $0.046 a minute; add the midpoints and you land near $0.13; add the tops and you reach $0.22. That spread, on the same functional product, is why per-minute pricing quoted without a breakdown tells you almost nothing.
Why is the all-in range so wide, from $0.05 to $0.40 a minute?
Two different things are being sold under one label. At the bottom of the range you are buying components and assembling them yourself. At the top you are buying an operated service, and most of the difference is the orchestration line, which is the only component whose band has no ceiling. A self-assembled stack lands at $0.05 to $0.15 a minute all-in. That number is real, and it is also incomplete, because it prices none of the engineering time to build it, none of the evaluation work to keep it from booking the wrong appointment, and none of the on-call burden when a carrier degrades at 4pm on a Friday. A premium managed service runs as high as $0.40 a minute. You are paying for someone else to own the integrations, the prompt regressions, the voice quality complaints and the pager. Whether that is expensive depends entirely on whether you were realistically going to do that work yourself. For a normal two-to-three-minute booking call, both models land at roughly $0.30 to $0.50 per completed call once everything is counted. That is usually the number worth arguing about, because it is the one that compares cleanly to what the call is worth.
How does an AI receptionist compare to a human receptionist?
This calculator will not tell you what a receptionist costs, because we cannot publish a salary figure that is true across markets, and inventing one would be exactly the kind of number this site has been cleaned of. Instead you type in your own fully loaded hourly cost, and the model does the comparison on your figure. Fully loaded means more than the wage. Include payroll taxes, benefits, paid time off, recruitment amortised over expected tenure, the desk, the software seat, and the supervisor's time. A common mistake is comparing a per-minute agent rate to a bare hourly wage, which flatters the agent by a wide margin. The comparison also has to be honest about coverage. A person covers the hours you staff. An agent covers all of them, and the fair comparison for after-hours volume is usually against an answering service or against nothing at all, not against your existing front desk. The other half the arithmetic misses: a person handles the walk-in, calms the angry customer, and notices that the same caller has phoned three times this week. An agent does not. The realistic deployment for most businesses is overflow and after-hours rather than replacement.
How do the $49 to $95 self-serve plans fit into this?
The self-serve market publishes flat monthly entry prices: Rosie at $49, Goodcall at $59, myaifrontdesk at $65, and Smith.ai at $95. Those are real list prices and they are genuinely cheap for low volume. The thing to check before comparing them to a per-minute build is what the entry tier includes. Flat plans cap minutes or calls, and the overage rate is where the economics actually live. At low volume a flat plan is almost always cheaper than assembling anything. At high volume the per-minute maths reasserts itself, and the crossover point is worth calculating rather than assuming. The second thing to check is what the agent is allowed to do. Answering and taking a message is a different product from booking into your calendar, checking availability, taking a deposit and writing the job back to your CRM. Cost per minute is only comparable between agents doing comparable work.
What costs does a per-minute price leave out?
Per-minute pricing is a usage rate. It is not the cost of the deployment, and treating it as one is the most common budgeting error in this category. The costs that sit outside the meter are the ones that decide whether the project succeeds: mapping what the agent is allowed to say and do, connecting it to the calendar and the system of record, building an evaluation set so you can tell whether a prompt change made things worse, recording and reviewing real calls, and the escalation path for when the agent should hand off to a person. For a production deployment, that build work is usually larger than the first year of usage cost. It is also the work that most distinguishes a deployment that quietly books jobs from one that gets switched off in month two.
What costs does a per-minute price leave out? — in detail
Discovery: what the agent answers, what it refuses, what it escalates
Integration: calendar, CRM, dispatch, payment, and the failure modes of each
Evaluation: a scored test set so prompt changes are measurable, not vibes
Call review: sampling real transcripts weekly for the first month
Number provisioning, porting, and compliance for recording and disclosure
What is a realistic all-in cost per minute for a voice AI agent in 2026?
Observed bands put a self-assembled stack at $0.05 to $0.15 per minute all-in, and a premium managed service at up to $0.40. For a normal two-to-three-minute call that is roughly $0.30 to $0.50 per completed call.
Which component costs the most?
Text-to-speech and the language model are the largest fixed contributors, at $0.015 to $0.030 and $0.010 to $0.060 per minute respectively. But orchestration and margin has the widest band, $0.010 to $0.100 and upward, and it is what separates a cheap stack from an expensive one.
Is an AI receptionist cheaper than hiring someone?
Usually yes on a pure per-minute basis, but the comparison only holds if you use a fully loaded staff cost including taxes, benefits, downtime and supervision, and only if you are honest that an agent does not do the non-phone half of the job.
Why do the $49 to $95 plans look so much cheaper than a per-minute build?
Because they cap volume. Flat entry tiers are genuinely cheaper at low call counts. Find the included-minutes limit and the overage rate, then use this calculator at your actual volume to see where the crossover sits.