Zillow lead qualification, showing scheduling, transaction coordination, and review collection. The agent that finally pays its own commission.
Real estate teams buy a lot of leads. Most of them go unworked because the SOI is slammed. AI is the only realistic path to a 100 percent contact rate on new leads in under 5 minutes. Below: the full stack from Zillow lead in to closing day.
Looking back at the last six deployments in this category, three things we would do differently:
Start the eval harness on day zero. We have always said this and we have always slipped it. The first time we shipped without a regression eval, we caught a prompt change that silently degraded conversion by 11 percent for two weeks before anyone noticed. Now we treat the eval harness as the first deliverable, before the first prompt.
Get the executive sponsor in the first user-acceptance session. Not the project manager, not the ops lead. The owner or the C-suite person whose name is on the budget. Their reaction to the first live test changes the trajectory of the project. Their feedback in week four is too late.
Document the human escalation paths before the AI ships. Every project we have shipped that did not have written escalation procedures had a moment in week two when an unexpected case hit, the AI escalated, and nobody knew who was supposed to handle it. Documenting the human side before the AI ships is half a day of work that prevents a week of fire-fighting.
If you want to talk through how any of this applies to your specific situation, grab a 20-minute call. We do not pitch on the call. If you would rather read more first, the docs and our comparisons cover most of the underlying technology choices in writing.
The most common failure mode with real-estate for vertical businesses is treating the problem as a model selection problem. It is not. The model is the easy part. The hard parts are the data pipeline feeding it, the eval that catches regressions, and the human ownership layer that keeps the system honest after the implementer leaves the building.
We have shipped this category of system enough times to recognize a few patterns. The teams that win allocate roughly 20 percent of project time to the model and prompts, 40 percent to data and integrations, 25 percent to evals and observability, and 15 percent to change management. The teams that lose flip those numbers, spend 70 percent on prompts, and end up with a great demo that nobody trusts.
The good news is that none of this is novel engineering. The patterns are well-understood now. The discipline to follow them is the rare part.
Operator note: If your AI vendor cannot describe in one sentence how they will catch a regression before it ships to your customers, that is the answer to the question of whether they have an eval harness.
Lead intake (voice + chat).
Showing day.
Transaction.
Closing.
ROI: a 20-agent team measured a 31% lift in lead-to-appointment conversion and recovered ~6 admin hours per agent per week. Total stack: $1,400/mo. Payback in week 3 from the lead-conversion lift alone.
Below is the production stack we ship for this category. It is opinionated. Other stacks work. This one ships in 72 hours and survives a real-world Monday morning.
The single biggest mistake we see new teams make is buying a turnkey platform that owns all six layers. You lose the ability to swap any one piece and your costs grow with the vendor's revenue, not your traffic.
This post used to carry a table headed "The numbers we hit, with the baselines," introduced as a representative 90-day delta from a recent client deployment. No deployment produced those numbers. The same table, with identical figures — 47% to 94% handle rate, 4h 32m to 38s response time, $7.10 to $1.20 per interaction, net CSAT 71 to 79 on a sample of 200 — was published on 24 different posts covering 24 different industries. Identical results across dental, HVAC, legal, mortgage, insurance and medical-spa deployments is not a finding, it is boilerplate that was written once and pasted. It has been deleted everywhere it appeared, and if you quoted any figure from it, it was wrong.
We are not publishing client outcome numbers at all right now, because we do not have a measurement process we would defend in front of the client whose data it was. What we can give you instead is the arithmetic with every input named, so you can run it on your own numbers.
| Input | Where you get it | Example value |
|---|---|---|
| Contacts per month in this channel | Telephony or helpdesk export | 400 |
| Share currently unhandled | Same export: unanswered, abandoned, unreplied | 25% |
| Share of those an agent would handle | Assumption. Start conservative | 60% |
| Close rate on handled contacts | Your CRM, trailing 90 days | 35% |
| Value of one closed outcome | Your CRM, trailing 90 days | $420 |
Recovered revenue per month = contacts x unhandled share x agent-handled share x close rate x outcome value. On the example inputs: 400 x 0.25 x 0.60 x 0.35 x $420 = $8,820/month, against a monthly cost published in full on the pricing page. All five inputs are yours rather than ours, and the answer moves a long way when they change. That is a model, and it is labelled as one.
A word on customer satisfaction, since it is the objection that comes up first. The conventional wisdom is that customers hate AI on the phone. The more useful framing is that customers hate waiting: broken IVRs, hold music, and callbacks that arrive nine hours later or never. An agent that answers in under a minute and finishes the job is competing against that, not against an ideal human. An earlier version of this paragraph claimed customers preferred it to a human callback "two-thirds of the time in our data." There was no such data and that figure has been deleted. Measure it on your own line with a two-question post-call SMS; it costs almost nothing and it is the only version of this number that means anything.
If you want help putting this into your business, book a 20-minute strategy call and we will sketch the stack on the call. Or run the numbers through our ROI calculator and see what the payback looks like for your shop.
We do not pitch on the call. If we are not the right fit, we will tell you and point you somewhere that is.