Vapi vs Retell vs Bland: which voice AI platform is actually cheaper?
Bland is cheapest and simplest if you want one all-in number: $0.14/min with no platform fee, falling to $0.11/min on a $499/mo plan. Vapi is cheapest at scale because its $0.05/min platform fee sits on top of at-cost models you can undercut with your own keys. Retell sits between them and is the easiest to reason about.
The options, and the losing case for each
Vapi — An orchestration layer that charges a thin platform fee and passes model costs through at cost. Best for: Teams that already have OpenAI, Deepgram, ElevenLabs or Cartesia contracts and want to bring their own keys. Vapi's platform fee is $0.05/min and model providers are billed at cost — $0 if you bring your own API key — so an organisation with negotiated model rates lands the lowest true cost per minute of the three. It is also the right pick when you need fine-grained control over the pipeline: swapping a transcriber mid-call, custom LLM endpoints, tool-calling into your own services. Not for: Anyone who wants a single predictable number. Vapi's per-minute cost is a sum of four or five line items that each move independently, and your invoice is not a per-minute rate — it is a platform fee plus concurrency lines plus whatever your model vendors charged. Concurrency past the included ten lines is $10 per line per month, which surprises teams that model volume but not simultaneity. And the compliance add-ons are priced like enterprise software: HIPAA is $2,000/mo and zero data retention is $1,000/mo, which makes Vapi a poor fit for a small regulated practice.
Retell AI — A voice platform that publishes every component's per-minute rate so you can build the number yourself. Best for: Teams who want the pipeline itemised but not assembled by hand. Retell publishes $0.055/min for its own voice infrastructure, $0.015/min for most TTS voices (ElevenLabs is $0.040/min), a per-model LLM rate, and $0.015/min telephony — and then quotes the assembled range as $0.07–$0.31/min. Twenty concurrent calls and ten knowledge bases are included at no charge, which is the most generous free concurrency of the three and matters enormously for inbound. Not for: Anyone whose cost model has to survive a model swap. Because the LLM is a separate per-minute line, moving from a cheap model to a frontier one can more than double your all-in rate — the published LLM lines run from $0.045/min to $0.16/min depending on the model. If your product team keeps upgrading the brain, your finance team keeps re-forecasting. Retell is also the least opinionated of the three about conversation design; you will write more of the state machine yourself than you expect.
Bland AI — A vertically integrated stack that quotes one all-in per-minute rate covering LLM, STT and TTS. Best for: Anyone who needs to put a defensible cost-per-call in a spreadsheet this week. Bland's rate covers the LLM, speech-to-text and text-to-speech with, in their words, no token charges and no model-provider pass-throughs. The free Start tier has no platform fee and no card requirement at $0.14/min, and the ladder to $0.11/min at $499/mo is legible. It is also the only one of the three that prices transfer minutes separately and cheaply, which matters if your agent's job is mostly to qualify and hand off. Not for: Teams who want to choose their own models. The all-in rate is only all-in because Bland picked the stack; if your use case needs a specific frontier model, a specific cloned voice, or a specific low-latency transcriber, the integration that makes Bland cheap is the thing standing in your way. The daily call caps also bite: 100 calls/day on Start and 2,000/day on the $299/mo Build plan are hard ceilings, not soft ones, and an outbound campaign will hit them.
Published pricing and platform behaviour, read off each vendor's own pricing page on 23 August 2026.
Platform fee — Vapi: $0.05/min (Build plan). Scale is an annual contract at volume pricing.; Retell AI: None. Pay-as-you-go only; components are billed individually.; Bland AI: $0 on Start · $299/mo on Build · $499/mo on Scale
All-in talk-time rate — Vapi: Platform fee plus model costs at cost ($0 with your own keys); Retell AI: $0.07–$0.31/min quoted range for voice agents; Bland AI: $0.14/min Start · $0.12/min Build · $0.11/min Scale
What the rate includes — Vapi: Orchestration only. STT, LLM and TTS billed at cost.; Retell AI: Voice infra $0.055/min + TTS $0.015/min (ElevenLabs $0.040) + LLM + telephony $0.015/min; Bland AI: LLM, STT and TTS. Bland states there are no token charges or model pass-throughs.
Telephony — Vapi: Billed at cost through your carrier or theirs; Retell AI: $0.015/min for US and most listed countries; Retell numbers $2.00/mo; Bland AI: Included in the talk-time rate; transfer minutes $0.03–$0.05/min
Included concurrency — Vapi: 10 lines, then $10 per line per month; Retell AI: 20 concurrent calls free, then $8 per concurrency per month; Bland AI: 10 on Start · 50 on Build · 100 on Scale
Daily call ceiling — Vapi: Not published as a hard cap; Retell AI: Not published as a hard cap; Bland AI: 100/day Start · 2,000/day Build · 5,000/day Scale · unlimited Enterprise
Free credit to test — Vapi: Free minutes on signup; the exact grant is not published as a figure; Retell AI: $10 in free credits, full platform access; Bland AI: No card required on Start
HIPAA / data controls — Vapi: HIPAA $2,000/mo · zero data retention $1,000/mo; Retell AI: Not published as a line-item price on the pricing page; Bland AI: Not published as a line-item price on the pricing page
Data retention (default) — Vapi: 14 days call history, 30 days chat on Build; custom on Scale; Retell AI: Not published on the pricing page; Bland AI: Not published on the pricing page
Build effort to first live call — Vapi: Highest — you assemble the pipeline and own the vendor relationships; Retell AI: Medium — components are pre-wired, conversation logic is yours; Bland AI: Lowest — one vendor, one rate, one console
Footnotes on that table
Latency and barge-in are deliberately absent from the price table. All three ship interruption handling and all three publish sub-second targets, but none of them publishes a measured, methodology-backed latency figure you could hold them to. We will not put a number in a table that the vendor has not committed to.
Vapi's Scale tier and Bland's Enterprise tier are both contracted to volume and are not published. If you are past roughly 100,000 minutes a month, none of the numbers above are the ones you will pay.
Why does a per-minute comparison of these three keep coming out wrong?
Because only one of the three is quoting the same thing. Bland quotes an all-in talk-time rate: $0.14/min on the free Start tier covers the language model, the transcription and the speech synthesis. Vapi quotes an orchestration fee: $0.05/min buys you the platform and nothing else, with model providers billed at cost and at $0 if you bring your own API keys. Retell quotes components: $0.055/min for its voice infrastructure, $0.015/min for most text-to-speech voices, a separate per-minute LLM line, and $0.015/min for telephony. Put those three numbers in a column and Vapi looks two-thirds cheaper than Bland. It is not. It is quoting a third of the stack. The only honest comparison is to assemble a full minute for each vendor at the same quality bar and then compare the assembled numbers — which is what the table above does, and why the Vapi column is a formula rather than a figure. The second reason the comparison goes wrong is concurrency. Voice is the one workload where volume and simultaneity are different variables. Ten thousand minutes spread evenly over a month is a completely different infrastructure bill from ten thousand minutes that all land between 8am and 10am on a Monday. Retell includes twenty concurrent calls free and charges $8 per additional concurrency per month. Vapi includes ten and charges $10 per line per month. Bland bundles concurrency into the plan tier. For an inbound receptionist workload with a sharp morning peak, that difference can exceed the difference in per-minute rate.
What does a realistic 3,000-minute month actually cost on each one?
Take a mid-sized inbound workload: 3,000 talk minutes a month, peak concurrency of about fifteen simultaneous calls, a mid-tier language model, and a standard synthetic voice rather than a cloned ElevenLabs one. On Bland's Build plan that is a $299 platform fee plus 3,000 minutes at $0.12, so $299 + $360 = $659, and the fifty included concurrent calls cover the peak with room. On the free Start tier the same 3,000 minutes is $420 with no platform fee, but ten concurrent calls will not survive a fifteen-call peak, so Start is not really an option for this shape of workload. On Retell, take the published components at the cheaper end: $0.055/min infrastructure, $0.015/min TTS, $0.015/min telephony, and a mid-tier LLM line. Retell's own quoted envelope for voice agents is $0.07–$0.31/min, and this configuration sits low in it. Twenty concurrent calls are included free, which covers the peak at no extra charge. Retell's phone numbers are $2.00/mo each. On Vapi, the platform fee alone is 3,000 × $0.05 = $150. Then add your own model bill. If you are bringing your own keys and buying tokens at your own negotiated rates, Vapi is very likely the cheapest of the three at this volume — and if you are not, you are paying at-cost retail for four vendors and Vapi's total will land close to Retell's. The five extra concurrency lines above the included ten cost $50/mo. The honest conclusion is that at 3,000 minutes these three are within a few hundred dollars of each other, and the choice should be made on build effort, model control and compliance rather than on rate. The rate only decides it above roughly 20,000 minutes a month, at which point Vapi with your own keys pulls clearly ahead and Bland's Enterprise contract becomes the thing to negotiate.
How do latency and barge-in actually differ between them?
Less than the marketing suggests, and not in a way any of them will put a number on. All three run the same fundamental loop — streaming transcription, an interruptible model call, streaming synthesis — and all three ship endpointing and barge-in. What differs is where the loop runs and how much of it you control. Bland's integration is the reason its rate is all-in: it owns the whole path, which means the latency budget is theirs to optimise and yours to accept. Vapi's is the opposite: because you can point it at your own transcriber, your own model endpoint and your own voice provider, your latency is mostly a function of choices you made, and a badly chosen frontier model will add a second of first-token delay that has nothing to do with Vapi. Retell sits in the middle, with a curated set of engines and published per-engine pricing that lets you trade cost against speed explicitly — the $0.015/min voices and the $0.040/min ElevenLabs voices do not behave identically. Barge-in is the thing to actually test, and it is not testable from a pricing page. The failure mode that kills real deployments is not that the agent cannot be interrupted; it is that it treats a caller's 'mm-hm' as an interruption and stops mid-sentence, or that it keeps talking over a caller who is trying to give a street address. Every one of these platforms exposes endpointing sensitivity as a tunable. Budget a week of real calls to tune it on whichever you pick, on all three if you can, and treat any vendor that tells you it works out of the box as not having run enough calls.
How do latency and barge-in actually differ between them? — specifics
Test barge-in with a caller giving a long alphanumeric string — an order number, a policy number, a licence plate. That is where endpointing breaks.
Test with background noise. A contractor calling from a job site is the median caller for most SMB deployments, not a quiet office.
Test a warm transfer to a human, because that is the moment your telephony configuration fails and no pricing page mentions it.
What about telephony — do you need Twilio underneath?
Not necessarily, and the answer changes the maths. Retell publishes $0.015/min for telephony in the US and most listed countries and rents numbers at $2.00/mo. Bland folds telephony into the talk-time rate entirely. Vapi lets you bring your own carrier. For reference, Twilio's own published US voice rates are $0.0085/min inbound on a local number, $0.0140/min outbound local, and $1.15/mo for a local number. So Retell's $0.015/min telephony line is roughly retail-plus, which is normal for a bundled resale and is not worth the integration effort to undercut at small volume. At large volume it is: a million minutes a year at a half-cent premium is real money, and that is exactly the point at which bringing your own carrier into Vapi starts to pay for the work. The thing to check before you sign anything is not the per-minute rate but whether the platform can port your existing business number, and whether it supports the STIR/SHAKEN attestation you need to avoid being labelled spam on outbound. A cheap minute that gets your calls flagged is not cheap.
Which one has the lowest total build effort?
Bland, clearly, and it is not close — provided your requirements fit inside the choices Bland has already made. One vendor, one console, one rate, one support relationship. If your agent needs to answer the phone, follow a script, look something up and book something, you will get there fastest here. Retell is the middle path and, for most teams building a real product rather than a single agent, the best default. The components are pre-wired but visible: you can see exactly what each part costs and swap the expensive ones. Twenty free concurrent calls means you can pilot without a capacity conversation. Vapi is the most work and the most control. Bringing your own keys means owning four vendor relationships, four sets of rate limits and four failure modes, and it means your on-call engineer needs to know which of them is down at 2am. That is a real, recurring cost that does not appear in any per-minute table. It buys you the ability to swap any component without leaving the platform, which is worth a great deal if you expect the model landscape to keep moving — and it has kept moving. A useful heuristic: if the person who will maintain this in six months is a founder, pick Bland. If it is a product engineer, pick Retell. If it is a platform team with existing model contracts, pick Vapi.
Figures we deliberately did not publish
Measured latency figures for all three. None of the three publishes a methodology-backed latency number on a first-party page, so we did not publish one either. Treat any third-party latency league table for voice AI as marketing until it shows you its harness.
Vapi's free-trial credit. The pricing page indicates free minutes on signup but does not state a dollar or minute figure we could quote.
HIPAA pricing for Retell and Bland. Neither publishes a line-item price; both route it through sales.
Price log — what was read, and where
Vapi: $0.05/min platform fee (Build); 10 concurrency included, $10/line/mo; HIPAA $2,000/mo; ZDR $1,000/mo; models at cost — read from https://vapi.ai/pricing on 2026-09-05
Retell AI: $0.07–$0.31/min voice agents; infra $0.055/min; TTS $0.015/min ($0.040 ElevenLabs); telephony $0.015/min; 20 concurrency free then $8/mo; numbers $2.00/mo; $10 free credit — read from https://www.retellai.com/pricing on 2026-09-05
Bland AI: $0.14/min Start (no platform fee) · $0.12/min + $299/mo Build · $0.11/min + $499/mo Scale; transfers $0.03–$0.05/min; 100/2,000/5,000 calls per day — read from https://www.bland.ai/pricing on 2026-09-05
Twilio: $0.0085/min inbound local · $0.0140/min outbound local · $1.15/mo local number — read from https://www.twilio.com/en-us/voice/pricing/us on 2026-09-05. Used only as the reference point for judging bundled telephony markups.
Is Vapi cheaper than Bland?
Only if you bring your own model API keys. Vapi's $0.05/min platform fee is not comparable to Bland's $0.14/min because Bland's rate includes the language model, transcription and speech synthesis and Vapi's does not. Assemble a full minute on Vapi at retail model rates and the two land close together; assemble it on negotiated model rates and Vapi wins clearly.
Which of the three is easiest to get live on this week?
Bland. One vendor, one rate, no card required on the Start tier, and the language model, transcription and speech synthesis are all included in the per-minute price. The trade is that you cannot choose the models. If the requirement is 'a working agent by Friday' rather than 'the right platform for three years', that trade is usually correct.
How much concurrency do I actually need?
Take your busiest hour, not your monthly total. A business taking 40 calls in its peak hour with an average handle time of three minutes needs roughly two simultaneous lines, plus headroom for clustering — call it four to six. Retell includes 20 free, Vapi includes 10, Bland ties it to the plan tier. Most SMB inbound workloads never leave the included allowance.
Do any of them include telephony, or do I need Twilio?
Bland includes it in the talk-time rate. Retell resells it at $0.015/min with numbers at $2.00/mo. Vapi lets you bring your own carrier. For reference, Twilio's own published US rates are $0.0085/min inbound local and $1.15/mo per local number, so Retell's bundled rate carries a modest markup that is not worth avoiding below very large volumes.