Intercom moved to per-resolution. Zendesk is still per-MAU. What each model actually costs you at scale, and which one we recommend by company size.
Intercom moved to per-resolution pricing in 2024 and the market noticed. Zendesk is still mostly per-MAU. Drift is hybrid. Ada is per-resolution. Custom builds are per-token. At scale, the gap between the cheapest model and the most expensive can be 5x for the same workload. Here is how to think about it.
The build sprint below runs on a 72-hour clock. That is the engineering, not the engagement: our published promise is live in 21 days from kickoff, which wraps this sprint in scoping, evals, shadow mode and cutover. If you see "72 hours" and "21 days" on this site and wonder which is true, both are — one is the part where code gets written.
Hour 0-8. Kickoff. Interview the two people who do this job today. Pull 50 sample inputs (calls, chats, tickets). Establish baseline metrics. Identify the three top customer intents.
Hour 8-24. First-pass prompt. Wire the orchestration. Stand up the eval harness with 25 cases drawn from the sample inputs. The eval harness has to exist before the first prompt does.
Hour 24-40. Integrations. CRM webhook, calendar booking, payment link if relevant. Each integration ships with a synchronous confirmation path.
Hour 40-56. Internal QA. The two people we interviewed in hour 0 spend 90 minutes running the agent through their hardest scenarios. Their feedback drives the second-pass prompt.
Hour 56-68. Shadow traffic. Real customer interactions, AI answers, human reviews before the answer is sent. We are looking for any case where the AI's draft is worse than the human's draft.
Hour 68-72. Cutover. We flip the routing rule, monitor for the first hour, and hand off the on-call rotation to the client's champion. The implementer stays on standby for 7 days.
Operator note: The 72-hour clock is real but it assumes the client has decided on success criteria before we start. If success criteria are unclear at hour 0, the clock does not start until they are. This is the single biggest cause of pilot drift we see.
Most teams measure too many things and then measure nothing. The 30-day measurement plan is short:
Five metrics, one dashboard, one Monday review. If the metric is not on the dashboard, it does not exist for the first 30 days.
After 30 days you can add CSAT, conversion-to-revenue, and channel attribution. Adding them earlier just adds noise.
At 10K interactions/month, the gap between pricing models is small. At 100K interactions/month, it matters a lot.
| Model | At 10K/mo | At 100K/mo | At 500K/mo |
|---|---|---|---|
| Per MAU (Zendesk-like) | $1,500 | $10,000+ | $35,000+ |
| Per message (Drift hybrid) | $1,200 | $11,000 | $52,000 |
| Per resolution (Intercom Fin) | $990 | $9,900 | $49,500 |
| Per token (custom build) | $200 | $1,800 | $8,500 |
The custom build looks like the obvious winner. It is not, because you are then on the hook for the platform: integrations, UI, analytics, identity, on-call. Add a half-time engineer ($75K/yr fully loaded) and the custom build at 100K/mo crosses break-even against Fin around month 9.
Our recommendation: per-resolution from a vendor up to ~200K interactions/month, custom build above that. The break-even is sensitive to your engineering cost.
Looking back at the last six deployments in this category, three things we would do differently:
Start the eval harness on day zero. We have always said this and we have always slipped it. The first time we shipped without a regression eval, we caught a prompt change that silently degraded conversion by 11 percent for two weeks before anyone noticed. Now we treat the eval harness as the first deliverable, before the first prompt.
Get the executive sponsor in the first user-acceptance session. Not the project manager, not the ops lead. The owner or the C-suite person whose name is on the budget. Their reaction to the first live test changes the trajectory of the project. Their feedback in week four is too late.
Document the human escalation paths before the AI ships. Every project we have shipped that did not have written escalation procedures had a moment in week two when an unexpected case hit, the AI escalated, and nobody knew who was supposed to handle it. Documenting the human side before the AI ships is half a day of work that prevents a week of fire-fighting.
If you want to talk through how any of this applies to your specific situation, grab a 20-minute call. We do not pitch on the call. If you would rather read more first, the docs and our comparisons cover most of the underlying technology choices in writing.
The most common failure mode with chat for pricing businesses is treating the problem as a model selection problem. It is not. The model is the easy part. The hard parts are the data pipeline feeding it, the eval that catches regressions, and the human ownership layer that keeps the system honest after the implementer leaves the building.
We have shipped this category of system enough times to recognize a few patterns. The teams that win allocate roughly 20 percent of project time to the model and prompts, 40 percent to data and integrations, 25 percent to evals and observability, and 15 percent to change management. The teams that lose flip those numbers, spend 70 percent on prompts, and end up with a great demo that nobody trusts.
The good news is that none of this is novel engineering. The patterns are well-understood now. The discipline to follow them is the rare part.
Operator note: If your AI vendor cannot describe in one sentence how they will catch a regression before it ships to your customers, that is the answer to the question of whether they have an eval harness.
If you want to pressure-test the numbers above against your business, book a call. It is free and we do not bring slides.
Prefer to read more first? Our case studies walk through three full deployments: what worked, what we would do differently, and what each one cost.