Order status and WISMO 20% to 40% of queue | Eliminate first, then self-serve, then automate the exceptions | Proactive delay notifications and accurate tracking kill most of it. A portal lookup takes the next slice. AI handles what is left: the stuck, late, and lost parcels that need a carrier read plus a judgment call. Fails when you jump straight to automating it, because the contact still gets created and, on a per-resolution meter, still gets billed. |
Returns and exchanges 10% to 20% | Self-serve portal, automate the policy exceptions | A portal with a policy engine and label generation handles the clean cases. Fails when the portal cannot express your real policy, so every edge case becomes a ticket anyway and you have paid for a portal that raised the average complexity of your queue. |
Refunds and store credit 5% to 15% | Automate, under a monetary ceiling | Write access to the order or payment system, a per-action cap, a daily aggregate cap, and an audit log. Fails when the AI can only explain the refund policy, which means a human still does the work and you have automated the easy half of a two-step job. |
Cancellations and subscription changes 10% to 25% on subscription brands | Automate, with a save flow | Write access to the billing or subscription system so pause, skip, swap, frequency, and address changes actually execute. Fails when the integration is read-only: the AI answers correctly and the human still clicks the button. This is also the one category that is a revenue lever, not just a cost one. |
Delivery address edits 3% to 8% | Automate, time-boxed | A write to the order before the fulfilment cutoff, and a hard stop after it. Fails when there is no cutoff logic, so the AI cheerfully promises a change on a parcel the warehouse already shipped. |
Product, sizing, compatibility 10% to 20% | Automate | Live catalog and policy grounding, your brand voice, and an explicit behaviour when the answer is not in scope. Fails when the catalog is not connected and the model fills the gap confidently. |
Damaged, defective, missing items 5% to 10% | Assist, then automate the clean cases | Evidence capture first, then a bounded replacement or credit for the unambiguous cases. Fails when you automate the judgment cases and take the CSAT hit on your most emotional contacts. |
Billing disputes and chargebacks 1% to 5% | Assist and escalate | The AI assembles the transaction history and evidence pack; a human decides. Fails when automated at all. |
Complaints, legal, safety, regulated claims 1% to 3% | Human only | Routing that recognises them on arrival, which is itself a good use of AI. Fails when the classifier is tuned for coverage instead of caution. |
VIP, wholesale, B2B accounts Varies | Assist | Full account context in one pane for a named owner. Fails when treated as routine volume. |