Customer support is the most-deployed agent vertical and the one with the clearest public evidence of both the upside and the failure cost. The load-bearing design question is not "can the model answer" — it usually can — but whose words legally bind you and what the agent is authorized to do to an account.
The two public benchmarks that frame the vertical
The upside: Klarna's AI assistant publicly reported handling two-thirds of support chats in its first month — 2.3M conversations, the workload equivalent of ~700 full-time agents, with matched satisfaction and a large drop in repeat inquiries. That scale is what makes deflection economics real.
The downside: in Moffatt v. Air Canada (2024), a tribunal held the airline liable for a bereavement-fare policy its website chatbot invented, rejecting the argument that the chatbot was a "separate legal entity responsible for its own actions." The precedent is blunt: your agent's answers are your company's representations. An invented policy is not a UX bug; it is a binding commitment made at scale.
Together these define the operating envelope: deflect aggressively where answers are grounded and reversible, and treat policy statements and account mutations as controlled actions.
Grounding: answer from documents, not from weights
The Air Canada failure was ungrounded generation. The fix is architectural, not prompt-level: the agent answers from versioned help-center content and cites the article it used, so every policy claim has a source that support leadership actually controls. No retrieval hit → say so and escalate; never fall back to the model's prior. Version the content store, and log which document version grounded each answer — when policy changes, you can identify what the agent told customers under the old version.
Authorization: intent × risk routing
Route every conversation on two axes. Intent determines which tools are even reachable (a shipping-status intent never needs the refund tool). Risk determines who decides: read-only lookups auto-resolve; state changes on an account require authenticated identity (authenticate before account access, not after the model has already read the record); money movement gets the same shape as every other high-stakes vertical — the model gathers evidence and explains eligibility, a deterministic policy service computes the amount, and exceptions cross a human approval gate. τ-bench — built specifically on airline/retail support scenarios with policy constraints — shows why: agents violate stated policy under multi-turn pressure at rates a single-turn demo never reveals, and its pass^k consistency metric collapses exactly on the tasks where authority matters.
Escalation is a feature, not a failure
Define hard triggers, and log them as first-class outcomes: low retrieval confidence, repeated tool failure, abuse, explicit human request, anything outside the tool allowlist. Two metrics keep the system honest over time: escalation precision (were the escalated cases genuinely hard?) and containment regret (of the auto-resolved cases, how many reopened or churned?). Deflection rate alone rewards confidently wrong answers — the exact behavior the tribunal priced.
Sources: Klarna AI assistant press release · Moffatt v. Air Canada, 2024 BCCRT 149 · τ-bench (arXiv:2406.12045).
Related: rag basics, human approval gates, guardrails & safety, agent evals, back-office operations agents, groundedness and hallucination.