The 72-hour runbook we use for new client deploys. Hour-by-hour, with the eval gates, rollback triggers, and the Slack handoff template.
We have shipped enough client agent deploys that we run them off a single 72-hour runbook. Hour by hour. Below: the full runbook, the eval gates between phases, and the rollback triggers.
Below is the production stack we ship for this category. It is opinionated. Other stacks work. This one ships in 72 hours and survives a real-world Monday morning.
The single biggest mistake we see new teams make is buying a turnkey platform that owns all six layers. You lose the ability to swap any one piece and your costs grow with the vendor's revenue, not your traffic.
Five failure modes show up in this category over and over. Each has a specific fix.
Drift in prompt voice. A prompt that worked in week one starts producing off-brand replies by week four because the model behind it silently versioned. Fix: pin the model version, run a weekly voice-drift eval against 50 canonical scenarios, alert on a 3-point deviation.
Stale retrieval. The KB updated, the embeddings did not. The agent confidently quotes last quarter's pricing. Fix: a freshness-check job that compares KB modified timestamps against embedding job runs hourly, and a hard ceiling that prevents serving any answer grounded in a document older than the freshness window.
Quiet hallucinations. The agent invents a policy or a part number with high confidence. Fix: every customer-facing answer must cite at least one source from retrieval, and the eval set includes 25 adversarial questions designed to bait hallucinations. No source, no answer.
Escalation breakdown. The agent escalates to a human, but the handoff context is one sentence and the customer has to start over. Fix: structured handoff payload (intent + history + sentiment + suggested next action), and a human eval pass on every 50th handoff.
Silent integration failure. The CRM webhook 500s, the agent acts as if it succeeded, the customer thinks the appointment is booked. Fix: synchronous confirmation back to the agent before it tells the customer anything is done, and a retry queue with paging on persistent failures.
Hour 0-4. Kickoff. Interview 2 ops people. Pull 50 sample inputs. Lock success criteria.
Hour 4-12. First prompt + eval harness scaffolding. 25 canonical test cases drawn from the samples.
Hour 12-24. Integrations. CRM, calendar, SMS. Each with synchronous confirmation back to agent.
Hour 24-36. Internal QA. The 2 ops people exercise the agent against their hardest scenarios. Feedback loop on prompts.
Hour 36-48. Shadow traffic. Real customer messages, AI draft, human review-and-send. Compare quality on 50 interactions.
Hour 48-60. 25% routing live. Monitor every interaction. Iterate prompts on visible misses.
Hour 60-68. 75% routing if metrics held. Hand off Slack + dashboards to the client champion.
Hour 68-72. Cutover to 100%. Implementer on standby for 7 days. First week is daily check-ins, then weekly.
Eval gates between phases:
If any gate fails, stop. Do not push through. The cost of a bad cutover is weeks of trust rebuilding.
The build sprint below runs on a 72-hour clock. That is the engineering, not the engagement: our published promise is live in 21 days from kickoff, which wraps this sprint in scoping, evals, shadow mode and cutover. If you see "72 hours" and "21 days" on this site and wonder which is true, both are — one is the part where code gets written.
Hour 0-8. Kickoff. Interview the two people who do this job today. Pull 50 sample inputs (calls, chats, tickets). Establish baseline metrics. Identify the three top customer intents.
Hour 8-24. First-pass prompt. Wire the orchestration. Stand up the eval harness with 25 cases drawn from the sample inputs. The eval harness has to exist before the first prompt does.
Hour 24-40. Integrations. CRM webhook, calendar booking, payment link if relevant. Each integration ships with a synchronous confirmation path.
Hour 40-56. Internal QA. The two people we interviewed in hour 0 spend 90 minutes running the agent through their hardest scenarios. Their feedback drives the second-pass prompt.
Hour 56-68. Shadow traffic. Real customer interactions, AI answers, human reviews before the answer is sent. We are looking for any case where the AI's draft is worse than the human's draft.
Hour 68-72. Cutover. We flip the routing rule, monitor for the first hour, and hand off the on-call rotation to the client's champion. The implementer stays on standby for 7 days.
Operator note: The 72-hour clock is real but it assumes the client has decided on success criteria before we start. If success criteria are unclear at hour 0, the clock does not start until they are. This is the single biggest cause of pilot drift we see.
Every architecture choice in this category is a tradeoff. Here are the ones we have made consciously, and the alternative we did not pick.
We use Pipecat instead of building our own orchestration. The win is months of saved engineering. The cost is being a release behind on a few model integrations.
We use Anthropic as the default LLM with OpenAI failover, not the other way around. The win is consistently lower hallucination rates in our eval set. The cost is slightly higher cost-per-token at the equivalent model tier.
We run a custom eval harness instead of using a vendor product. The win is the eval set lives in the same Git repo as the prompts, so PRs that change prompts fail the build if they break an eval. The cost is we own the upkeep.
We use Twilio over a cheaper telephony provider. The win is concurrency ceiling and the depth of the diagnostic tools. The cost is roughly 18 percent more per minute.
We deploy on Fly.io, not Vercel. The win is real persistent connections and global edge POPs that survive long-lived voice sessions. The cost is more devops overhead.
These tradeoffs are not laws. They are starting points. If you tell us you have an existing GCP estate, we will adapt the stack. The principles do not move; the implementations do.
This post is the short version. The long version takes 90 minutes and a whiteboard. If you want the long version, grab a slot and bring your worst metric. We will work backward from there.
Or, if you are still in the read-and-think phase, our docs and comparisons pages have most of the answers in writing.