CSAT, AHT, NPS. Ignore them all for the first 90 days. Deflection rate is the only number that tells you whether the chat agent is paying rent.
Every chat agent vendor will show you a dashboard with eight metrics. CSAT, AHT, NPS, sentiment, escalation rate, response time, abandonment, resolution time. Ignore all of them for the first 90 days. The only number that tells you whether the chat agent is paying rent is deflection rate, and most vendors define it wrong.
Five failure modes show up in this category over and over. Each has a specific fix.
Drift in prompt voice. A prompt that worked in week one starts producing off-brand replies by week four because the model behind it silently versioned. Fix: pin the model version, run a weekly voice-drift eval against 50 canonical scenarios, alert on a 3-point deviation.
Stale retrieval. The KB updated, the embeddings did not. The agent confidently quotes last quarter's pricing. Fix: a freshness-check job that compares KB modified timestamps against embedding job runs hourly, and a hard ceiling that prevents serving any answer grounded in a document older than the freshness window.
Quiet hallucinations. The agent invents a policy or a part number with high confidence. Fix: every customer-facing answer must cite at least one source from retrieval, and the eval set includes 25 adversarial questions designed to bait hallucinations. No source, no answer.
Escalation breakdown. The agent escalates to a human, but the handoff context is one sentence and the customer has to start over. Fix: structured handoff payload (intent + history + sentiment + suggested next action), and a human eval pass on every 50th handoff.
Silent integration failure. The CRM webhook 500s, the agent acts as if it succeeded, the customer thinks the appointment is booked. Fix: synchronous confirmation back to the agent before it tells the customer anything is done, and a retry queue with paging on persistent failures.
The build sprint below runs on a 72-hour clock. That is the engineering, not the engagement: our published promise is live in 21 days from kickoff, which wraps this sprint in scoping, evals, shadow mode and cutover. If you see "72 hours" and "21 days" on this site and wonder which is true, both are — one is the part where code gets written.
Hour 0-8. Kickoff. Interview the two people who do this job today. Pull 50 sample inputs (calls, chats, tickets). Establish baseline metrics. Identify the three top customer intents.
Hour 8-24. First-pass prompt. Wire the orchestration. Stand up the eval harness with 25 cases drawn from the sample inputs. The eval harness has to exist before the first prompt does.
Hour 24-40. Integrations. CRM webhook, calendar booking, payment link if relevant. Each integration ships with a synchronous confirmation path.
Hour 40-56. Internal QA. The two people we interviewed in hour 0 spend 90 minutes running the agent through their hardest scenarios. Their feedback drives the second-pass prompt.
Hour 56-68. Shadow traffic. Real customer interactions, AI answers, human reviews before the answer is sent. We are looking for any case where the AI's draft is worse than the human's draft.
Hour 68-72. Cutover. We flip the routing rule, monitor for the first hour, and hand off the on-call rotation to the client's champion. The implementer stays on standby for 7 days.
Operator note: The 72-hour clock is real but it assumes the client has decided on success criteria before we start. If success criteria are unclear at hour 0, the clock does not start until they are. This is the single biggest cause of pilot drift we see.
Most vendors define deflection rate as 'percent of conversations that did not escalate to a human'. That metric is gameable. If the agent says 'I cannot help with that, here is a help article link', the conversation did not escalate but the customer is also no closer to resolved.
Our definition: deflection rate = (resolved by agent) / (total contacts in channel). Resolved means the customer's stated goal was achieved AND no follow-up contact happened in the next 7 days.
Operationally we measure it like this:
WITH first_contacts AS (
SELECT conversation_id, customer_id, started_at, resolved_by_ai
FROM conversations
WHERE channel = 'web_chat' AND started_at >= NOW() - INTERVAL '30 days'
),
followups AS (
SELECT customer_id, MIN(started_at) AS next_contact
FROM conversations
WHERE channel IN ('web_chat','email','phone')
GROUP BY customer_id
)
SELECT
COUNT(*) AS total,
COUNT(*) FILTER (
WHERE resolved_by_ai AND (next_contact IS NULL OR next_contact > started_at + INTERVAL '7 days')
) AS deflected,
ROUND(100.0 * COUNT(*) FILTER (
WHERE resolved_by_ai AND (next_contact IS NULL OR next_contact > started_at + INTERVAL '7 days')
) / NULLIF(COUNT(*),0), 1) AS deflection_rate
FROM first_contacts f
LEFT JOIN followups u USING (customer_id);
Run this weekly. The 7-day window catches the 'sorta resolved' cases that show up later as escalations.
Most teams measure too many things and then measure nothing. The 30-day measurement plan is short:
Five metrics, one dashboard, one Monday review. If the metric is not on the dashboard, it does not exist for the first 30 days.
After 30 days you can add CSAT, conversion-to-revenue, and channel attribution. Adding them earlier just adds noise.
Looking back at the last six deployments in this category, three things we would do differently:
Start the eval harness on day zero. We have always said this and we have always slipped it. The first time we shipped without a regression eval, we caught a prompt change that silently degraded conversion by 11 percent for two weeks before anyone noticed. Now we treat the eval harness as the first deliverable, before the first prompt.
Get the executive sponsor in the first user-acceptance session. Not the project manager, not the ops lead. The owner or the C-suite person whose name is on the budget. Their reaction to the first live test changes the trajectory of the project. Their feedback in week four is too late.
Document the human escalation paths before the AI ships. Every project we have shipped that did not have written escalation procedures had a moment in week two when an unexpected case hit, the AI escalated, and nobody knew who was supposed to handle it. Documenting the human side before the AI ships is half a day of work that prevents a week of fire-fighting.
If you want to talk through how any of this applies to your specific situation, grab a 20-minute call. We do not pitch on the call. If you would rather read more first, the docs and our comparisons cover most of the underlying technology choices in writing.
If you want help putting this into your business, book a 20-minute strategy call and we will sketch the stack on the call. Or run the numbers through our ROI calculator and see what the payback looks like for your shop.
We do not pitch on the call. If we are not the right fit, we will tell you and point you somewhere that is.