FIELD NOTES

AI voice agents in 2026: how we route between Anthropic, OpenAI, Google and Groq

A production voice agent should not pick one LLM. Here is the call-state routing we actually ship, the failover chain behind it, a cost-per-call model built from published token pricing, and a correction retracting the benchmark this post used to carry.

If you read any AI vendor blog, every LLM is somehow the best at every task. That's marketing. We run AI receptionists in production for service businesses, and we have informed opinions on which models actually pick up the phone well — and which fall over the second a real customer interrupts.

The four models that matter for SMB receptionist work in mid-2026:

  • Anthropic Claude Sonnet 4.6 (the Claude family's voice-optimized variant)
  • OpenAI GPT-5 (the flagship reasoning model, low-latency variant)
  • Google Gemini 2.5 Pro (the multimodal model)
  • Groq-served Llama 3.3 70B (the fastest-served open-weights option)

The conclusion this post exists for: the best production voice agent in 2026 routes between models based on call state, rather than picking the smartest one and using it for everything. The TrainYourAgent runtime chain is Anthropic as primary, Groq as fast-fallback, and Gemini as final fallback — three providers, which you can read in api/_lib/llm.ts. Why we made each of those choices is below.

Correction: the four-model scorecard has been deleted

An earlier version of this post opened by claiming 600 controlled voice calls, 150 per model, placed by a 12-person panel of native and non-native English speakers against an identical Deepgram-plus-ElevenLabs-plus-Vapi stack, and reported a scorecard of median time-to-first-token, end-to-end latency, filler rate, interruption-recovery rate, accent-comprehension rate and on-call conversion for each of the four models, with a stated confidence interval.

No such test was run. There was no panel, no 600 calls, and no scorecard. Every number in that table was invented, and it has been deleted. So has the prose built on it: that Groq hit 287 ms and Anthropic 612 ms, that Gemini led accent comprehension at 92%, that Anthropic converted at 88% against Groq's 73%. If you quoted any of those figures, they were wrong and I am sorry.

I am not going to replace them with softer invented numbers either. What is left below is the part that was always real: an architecture argument, a routing table that matches shipped code, and a cost model derived from published per-token pricing that you can re-derive yourself.

Why route at all, instead of picking one model

The reason to route is that the four models have genuinely different shapes, and the differences are well attested in each vendor's own published material rather than in anything we measured:

  • Served open-weights models are built for speed. Groq's whole product is inference throughput, and its published position is time-to-first-token, not reasoning depth. That is the right property for the first two seconds of a call.
  • Frontier reasoning models are built for the hard turn. Anthropic and OpenAI both publish instruction-following and tool-use behaviour as the thing they optimise. That is the right property for a multi-objection booking or a complaint.
  • Google's models carry the heaviest multilingual weighting. Gemini's own documentation leads on multilingual and multimodal coverage. For a customer base with a meaningful non-native-English share, that matters.

None of that requires a benchmark from us. It requires reading what each vendor says it built, then putting each model where its shape fits. Whether the shape holds up on your traffic is something you have to test on your traffic, with your accents, your knowledge base and your call patterns — which is the same advice we give about latency, and for the same reason.

What we actually do in production

The TrainYourAgent stack does not pick a single LLM. We route between them based on call state. The reasoning:

  1. Greeting + intent capture: Groq Llama 3.3 70B. Speed matters most here. The customer is impatient and the conversation is shallow.
  2. Knowledge-base lookup + answering: Anthropic Sonnet 4.6. Quality matters most here. The model needs to give the right answer.
  3. Objection handling + close: Anthropic Sonnet 4.6 with a 1.5-second think-step. Closing a booking requires conversational nuance and the model needs to track multi-turn objections.
  4. Multilingual handling: Gemini 2.5 Pro on detected accent / language switch. Comprehension matters most here.
  5. Hard escalation reasoning (price negotiation, complex multi-service quote, complaint triage): OpenAI GPT-5 for its broader reasoning depth. Used in under 5% of calls but high-stakes.

This is the multi-LLM fallback story. We also use Anthropic-to-Groq-to-OpenAI as a hard failover chain — if Anthropic's API is degraded (which has happened twice in the last 90 days for periods of 4–14 minutes), the agent falls back to Groq and the call doesn't drop.

This is a real engineering decision, not a marketing claim. The case for multi-LLM is:

  • Quality (route the right model to the right turn)
  • Uptime (no single provider's outage kills your agent)
  • Cost (use the cheap model where it suffices)

The case against:

  • Engineering complexity
  • Subtle behavioral inconsistency across turns
  • Harder evals (you have to test the routing, not just each model)

For most operators, this is engineering work they don't want to own — which is the case for buying a multi-LLM-aware platform rather than building one.

What we don't recommend

A few things we tested and disliked:

  • Single-model OpenAI Realtime API. The end-to-end TTS is good but the latency variance is high (we observed sub-second responses and 4-second responses on identical traffic, within the same call). For a production SMB receptionist this variance is unacceptable.
  • Whisper for STT. Deepgram Nova-3 is meaningfully better at accented and noisy-environment calls. We don't ship Whisper to customers anymore.
  • Browser-based agents (Vapi-only without LLM control). The agent works for low-traffic demos and breaks at scale; we recommend Vapi as orchestration with explicit LLM routing under it.

Cost model (arithmetic, not a measurement)

This is a model with stated inputs, not something we observed. Assume a call consumes ~3,000 input tokens and ~800 output tokens, and multiply by each vendor's published per-token list price as of June 2026. Re-derive it yourself from the pricing pages linked at the bottom — the token counts are the only assumption, and they are the input to change if your calls run longer.

Model Cost / call
Anthropic Sonnet 4.6 $0.024
OpenAI GPT-5 $0.041
Google Gemini 2.5 Pro $0.029
Groq Llama 3.3 70B $0.004

For a 14-truck HVAC operator at 612 calls/month, the all-LLM bill ranges from $2.45/month (Groq-only) to $25.10/month (OpenAI-only). The LLM is rarely the dominant cost; STT, TTS, and telephony together typically run 4–8x the LLM line. Pick the right model for the right call state and the cost differential becomes noise.

What this post is not

  • It is not a benchmark. We have not run a controlled comparison of these four models on voice traffic, and we are not going to publish one until we have a harness whose results we would hand to a customer. Anyone telling you which LLM is fastest on the phone should be asked, in writing, what they measured and how.
  • The routing table is an architecture, not a proof. It says where we put each model and why. It does not prove that arrangement beats a single-model build on your traffic. If someone shows you a number that claims to, ask the same question.
  • The cost table is list-price arithmetic. It moves whenever a vendor reprices, and it ignores caching, batching and negotiated rates, all of which cut the real bill.
  • Cost-tier downgrades are not covered. Anthropic Haiku, OpenAI's mini tier and Gemini Flash can be excellent for receptionist work; we scope those per customer rather than generalising.

Sources and references

The bottom line

The "best LLM for voice" question has the same answer as the "best language for software" question: the right one for the layer. Use the fast model for the call-state where speed matters, the smart model where reasoning matters, and the specialist model where language matters. Single-model voice agents in 2026 are leaving real call quality on the table.

If you don't want to make these routing decisions yourself, that's what the TrainYourAgent platform handles. The model routing is invisible to the operator — your customer gets the right model on every turn without you needing to know which one. See how the platform handles it, or book a 20-minute scoping call.

Filed under