The prompt patterns that make a support agent safe on either model, the failure modes each one has, a list-price cost model, and a correction retracting the 60-ticket bake-off this post used to claim.
Pricing note, 5 September 2026. Our published prices changed on this date: Operators moved from $4,950 build / $1,997 per month to $7,500 / $2,950, and Agent in a Day moved from $497 one-time to $2,500 build / $497 per month. The arithmetic below was run against the prices in force when it was written and has been left as it was rather than quietly restated. Current prices are on the pricing page.
I get asked weekly which model to use for customer support. The "which is smarter" question is the wrong one. The right one is: which one ships the right answer in under 4 seconds, costs under $0.04 per resolved ticket, and doesn't hallucinate a refund policy that doesn't exist? That's a different bake-off.
An earlier version of this post said we ran one: "60 real anonymized support tickets from three TrainYourAgent customers (a DTC skincare brand, a SaaS onboarding team, a roofing scheduling desk)" through GPT-5 and Claude Sonnet 4.6 with an identical rubric and a blind third-model evaluator, split into three buckets of twenty. It then published a scorecard — factual accuracy 91.7% against 94.2%, empathy 7.8 against 8.5, hallucinated policy 7/60 against 2/60, p50 and p95 latency, cost per thousand tickets.
That bake-off was never run. No customer tickets were used, there was no evaluator, and every number in that scorecard was invented. It has been deleted, along with the counts built on it: 18/20 against 13/20 on the refund case, zero schema violations across sixty tickets, a 25% latency edge, a 70% cut in hallucinated policy, a 12% lift in resolution rate. If you quoted any of them, they were wrong and I am sorry. Using customers' support tickets as the setting for invented results was the worst part of it, and it should never have been written.
What survives is the part that was always real and is checkable by you in an afternoon: two prompt patterns, the failure modes each model actually has, and a cost model built from published list prices. The example replies below are illustrations written to show the shape of the difference — they are not transcripts, and they are not evidence.
Run your own comparison with this and you will have better data than anything a vendor blog can give you:
You are a support agent for [BRAND]. Use only the provided knowledge base
to answer. If the knowledge base does not contain the answer, say so plainly
and offer to escalate. Be warm, brief (under 6 sentences when possible),
and do not invent policy. End every reply with a single concrete next step.
Hold the rest constant when you compare: same retrieval (top-5 chunks from your KB, retrieval temperature 0.0), same model temperature, same rubric. Score on four things that actually decide whether a support agent is safe to ship — factual accuracy against the KB, appropriate empathy, brevity, and safe escalation behaviour. Then read fifty of the replies yourself. There is no substitute for that last step, and no scorecard from us replaces it.
If your support tickets are 60%+ "my integration broke, here's the stack trace," GPT-5's code reasoning is worth testing first. Both vendors publish coding-benchmark results and OpenAI's have historically led on traceback-style debugging, which is the closest public proxy for this ticket type. Test it on your own tracebacks before you commit — this is the bucket where the gap is most likely to be real and most likely to matter.
The small model tiers change the arithmetic at volume. At 50k tickets/month in a mixed-language ecommerce flow, the gap between a frontier tier and a nano tier is the difference between a few hundred dollars a month and a few tens. Work it out from current list prices rather than from a figure in a blog post — they move. The tradeoff to watch is that small tiers give up accuracy, so they belong on low-stakes tickets only, and you should measure the drop on your own ticket mix before routing anything important to them.
If your support agent has to emit {intent, urgency, sentiment, suggested_action, draft_reply, escalation_target} every single time and you parse it programmatically, schema adherence is the property to test, and it is the easiest one to test properly: run a few hundred tickets, validate every output against the schema, count the failures. Both vendors now offer constrained or structured output modes that make this close to a non-issue — use them rather than trusting either model's freehand JSON, and null-versus-empty-string is the mismatch to watch for.
The clearest difference in shape shows up on angry tickets. An illustrative ticket:
"I've been a customer for 4 years and this is the SECOND time you've billed me twice. I want my money back and I want to know why this keeps happening. This is unacceptable."
A typical GPT-5-shaped reply:
"I understand this is frustrating. Let me look into your account. I see two charges on March 15th for $89 each. I'll process a refund for the duplicate charge and it should appear in 5-7 business days. Is there anything else I can help with?"
A typical Claude-shaped reply:
"I'm really sorry — being double-billed twice in your fourth year as a customer is exactly the kind of thing that breaks trust, and I don't blame you for being frustrated. I can see both charges on your account from March 15th. I'm refunding the duplicate now and you'll see it back in 5-7 business days. Separately, I'm flagging this to our billing engineering team so we can figure out the root cause — I'll personally make sure someone follows up with you within 48 hours with what we found and what we changed. If you want to escalate to a manager in the meantime, just say the word."
The second reply lands more of the emotional beats and proposes a root-cause follow-up. Which model produces which shape on your prompt and your knowledge base is exactly the thing to check by reading replies, and it moves between model versions.
An illustrative ticket: "Can I get a refund? I bought the annual plan in January and want to cancel now."
The KB explicitly states: "Annual plans are non-refundable after the 14-day satisfaction window. Pro-rata credits may be issued at customer-success discretion."
A model that invents policy answers: "I can refund you $X for the remaining months." There is no such policy, and the customer now believes there is.
A model that respects the KB answers: "Looking at your account, you're past the 14-day window, so a full refund isn't standard. What I can do is offer a pro-rata credit toward a future purchase — that's at our team's discretion and I think your situation qualifies. Want me to apply it?"
Inventing a policy that does not exist is the single biggest support-agent failure mode, and it is the one to build your eval set around: write twenty questions whose correct answer is a refusal or a bounded offer, and count how often each model gets them right. Do not take a ranking from anyone, including us, without that.
An illustrative ticket: "I think I'm having a panic attack from the stress of this dispute. Please help me."
A model that misses the signal answers: "I can help you with your dispute. Can you tell me your account email?"
A model that catches it answers: "I want to make sure you're okay first. If you're feeling overwhelmed, please take care of yourself — and if you need someone to talk to immediately, you can reach the 988 Suicide & Crisis Lifeline by calling or texting 988. I'm here to help with the dispute whenever you're ready, and I'm also looping in a human teammate so you don't have to navigate this alone."
Missing a safety signal is rare and catastrophic, and it is the one behaviour to test adversarially on every model you consider and every prompt revision, before either touches a customer. Write ten tickets that carry a distress signal alongside an ordinary request, and check that the agent catches the signal, escalates, and still handles the request. Whichever model you pick, that test belongs in your regression suite rather than in a blog post.
We do not publish a latency comparison for these two models, because we have not measured one in a way we would defend. What is worth knowing is where the latency comes from, because it is not usually the model.
The dominant term on a reasoning-tier model is reasoning effort applied to tickets that did not need it. Both vendors now expose a control for this. If you are running an agent where every second hurts — a voice agent over a phone line, live chat with a five-second visible response budget — turn that control down before you switch models, and measure p50 and p95 on your own traffic. A configuration change frequently beats a model change, and it is reversible.
This is a model with stated inputs. Take 10,000 resolutions a month, assume a typical support turn of roughly 3,000 input and 500 output tokens plus a retrieval call, and multiply by each vendor's published list price. At current frontier-tier rates that lands both models in the low hundreds of dollars a month, within a few tens of dollars of each other, with retrieval adding a similar order again.
Set that against the human comparison on the same 10,000 contacts: outsourced support at a per-contact rate in the region of $0.85 is a five-figure monthly line. Re-derive both sides yourself — the inputs are your ticket volume, your token counts and today's list prices.
The delta between the two LLMs is small. The delta to humans is two orders of magnitude. The question is not "which model," it is "are you running an agent at all?"
This is what we ship by default and why. It is a starting configuration, not a ranking, and the reasons are architectural rather than measured:
Neither of these is scored here, because we have no scores. Both are cheap to test on your own queue, and both address a failure mode you can see with your own eyes inside an hour:
Pattern 1 — Cite-or-decline. Add to the system prompt:
For every factual claim about policy, prices, or timelines, cite the
source KB chunk in brackets like [doc: refund-policy.md]. If you cannot
cite a source, refuse to make the claim.
This is the single highest-leverage line in a support system prompt. Measure the effect on your own adversarial set rather than trusting a number — but expect it to be large, because it converts an open-ended generation problem into a retrieval problem the model can fail loudly instead of quietly.
Pattern 2 — Single concrete next step. Add:
End every reply with exactly one concrete next step in the format:
"Next: [I will / Please / We will] ..."
Customers reply faster when they know exactly what to do, and a reply with no next step is the most common reason a resolved-looking ticket bounces back. This one is cheap to A/B on your own queue.
If you want to grade your own support-agent prompts on the rubric described here, run them through our Prompt Critic. It scores on clarity, specificity, structure and safety, and returns a rewritten version of any prompt you submit.
Or book a call if you want us to ship the support agent for you. Published pricing: Self-Serve at $99/mo, Operators at a $4,950 build fee plus $1,997/mo with 5,000 minutes included, Scale at $9,950 plus $4,997/mo. Live in 21 days or the build fee is refunded. Full matrix on the pricing page.