Endless loops, hallucinated pricing, hangup-on-silence, robot voice, and the dreaded I am an AI. Here is the specific fix for each.
After 14 months of running voice agents in production, the same five failure modes show up over and over. They are not bugs in the model. They are gaps in how we wrap the model. Below: the five failure modes, what they sound like to the caller, and the specific engineering fix for each. What we would do differently next time Looking back at the last six deployments in this category, three things we would do differently: Start the eval harness on day zero.
We have always said this and we have always slipped it. The first time we shipped without a regression eval, we caught a prompt change that silently degraded conversion by 11 percent for two weeks before anyone noticed. Now we treat the eval harness as the first deliverable, before the first prompt. Get the executive sponsor in the first user-acceptance session. Not the project manager, not the ops lead.
Endless loops, hallucinated pricing, hangup-on-silence, robot voice, and the dreaded I am an AI. Here is the specific fix for each. It is filed under AI Voice because that is where operators looking for this problem actually start, and it is written from production work rather than from a content calendar.