FIELD NOTES

How We Reverse-Engineered Bland.ai's Pricing and Built a Better Stack for Half the Cost

Bland charges a flat $0.09/min and won't tell you why. We pulled apart their stack, replicated it on commodity components, and ended up at $0.041/min. Here's exactly how the cost build-up works, line by line.

Bland.ai is one of the cleanest packaged voice-AI products on the market. Single endpoint, predictable $0.09/min, batteries included. For a lot of customers that's exactly the right answer.

But $0.09/min on a million minutes is $90k. And once a client crosses that threshold, the math gets uncomfortable. So a few months ago we ran an experiment: could we replicate Bland's behavior — the latency, the natural turn-taking, the human-sounding voice — on a commodity stack we owned? And what would it actually cost?

Spoiler: yes, $0.041/min, with one important caveat we'll get to.

What Bland is actually doing

Bland doesn't expose its internals. What follows is inference from public material and from ordinary use of the API — their docs, their pricing page, response headers, and what the product visibly does. It is not the output of a measurement campaign. An earlier version of this section claimed we sent "~4,000 test calls" and analysed the audio and response latency to derive component-level timings, and quoted an end-to-end figure of 600-750ms for Bland alongside a 450ms figure for a Vapi-plus-Cartesia stack. No such campaign was run and those timings were invented. They have been deleted, and they were also flatly contradicted by the correction two sections below, which is how they were caught.

What can be said about the shape of the stack, without timings attached:

  • STT: a custom-tuned Whisper-family model rather than a third-party API, on their own inference.
  • LLM: their own fine-tune, optimised for short action-oriented replies. Bland describes this in their own materials; the base model is not published.
  • TTS: their own neural voice clone, roughly a dozen voices. Bland publishes no first-byte latency figure, so we are not quoting one.
  • Telephony: SIP into either Twilio or their own carrier interconnect — Twilio call SIDs appear in some response headers, which suggests a hybrid.
  • Orchestration: WebSocket-based, with interruption and barge-in handled at the audio layer.

We are not publishing an end-to-end latency figure for Bland, ours, or anyone else's, for the reason set out in the correction below.

The replication stack

We picked components that match each Bland piece on quality and beat it on price:

Bland piece Our replacement Cost / min
Custom Whisper STT Deepgram Nova-3 streaming $0.0043
Custom 70B fine-tune GPT-4.1-mini with a strict system prompt + 5-shot examples $0.003
Custom voice clone Cartesia Sonic English (cloned voice via 30s sample) $0.025
SIP / carrier Twilio Elastic SIP $0.0085
Orchestration Pipecat self-hosted on Fly.io (3 machines, $40/mo each) $0.0008 amortized
Total $0.0416 / min

That's 54% less than Bland's $0.09/min. On a million minutes that's $48,400/year saved. On 5M minutes it's a quarter-million dollars.

The latency question, and what we can't tell you

This section used to claim a measured median end-to-end latency for both Bland and our own stack, and asserted that ours was faster. Those numbers were not measured. They have been deleted, along with the component-level comparison that was derived from them. Everything above this line is a cost build-up from published per-unit rates, which you can re-derive yourself; the latency claim was not that, and it should not have been sitting next to work that was.

What we can say honestly:

  • We have not run a controlled latency comparison against Bland, or against any other vendor. We do not publish a median first-audio figure for our own stack either, because we have not measured one in a way we would defend.
  • We engineer to a target, not to a benchmark: under 800ms voice-to-voice. That is a design constraint we build against and it is what we hold ourselves to internally. It is not a measurement, and it is not a claim about anyone else's product.
  • The architectural argument still stands on its own, and it is the honest version of what this section was reaching for: every provider boundary is a place for latency to accumulate, so a self-assembled stack gives you the ability to tune — colocate in one region, pick a streaming-first TTS, keep the model small enough that first-token is quick. Whether you actually end up faster than a packaged product depends entirely on whether you do that work.

If latency is the deciding factor for you, the only number worth anything is the one you measure yourself: call both, from the phone and network your customers use, interrupt the agent mid-sentence, and call again after five idle minutes to catch cold start. Our comparison post on Bland vs Synthflow vs Vapi sets out what each vendor publishes about latency and why those figures are not comparable with each other.

The real trade in this rebuild was never latency. It was cost against operational surface area, which is the next section.

What we lost

Three things got worse, not better:

  1. Voice naturalness. Bland's voice clone has a slightly more human cadence than Cartesia, especially on long sentences — a subjective judgement from listening to both, not a preference study. An earlier version of this line claimed "~80% of testers couldn't tell the difference; 20% slightly preferred Bland." There were no testers and that split was invented. It has been deleted. We mitigated the gap by writing shorter sentences into the system prompt, which helps regardless of which TTS you land on. Judge the voices yourself on the two demo lines; it takes two minutes and it is the only assessment that counts.
  2. Operational surface area. We now own a Pipecat deployment, a Cartesia account, a Deepgram account, an OpenAI account, a Twilio account, and a metrics dashboard. That's five vendors and one piece of infra to monitor. Bland was one phone call to support.
  3. Failure modes are ours now. When Cartesia has a regional outage (it has happened), it's our pager that goes off. With Bland we'd just file a ticket and wait.

When you should NOT do this

If you're under 100k minutes/month, stay on Bland or Vapi. The savings ($300–500/month) don't pay for the engineering time. You will spend more in your own hours debugging Pipecat configs than you save in inference fees.

The break-even, in our experience, is around 250k minutes/month. Below that, the convenience tax is the right answer. Above that, the math demands you take ownership.

The real lesson

Reverse-engineering Bland wasn't really about Bland. It was about understanding that every "AI platform" right now is a thin orchestration layer over commodity APIs. The platform exists because most customers don't want to glue four vendors together. That's a real service and worth paying for. But the moment you have the volume to justify the engineering, the underlying components are sitting there — same APIs the platforms are calling, same prices.

That's the arbitrage. Use it when it's big enough to matter. Don't bother when it isn't.

Filed under