The test set is the asset. Not the vendor's dashboard.
The method, in seven steps
- Sample two sets from 90 days of your own resolved tickets. A volume set of about 200 cases drawn in proportion to intent frequency, and a risk set of 60 to 100 cases deliberately over-weighted toward money, policy exceptions, and anger.
- Write every case as an expected outcome, never an expected wording. The decision, the action to take or refuse, and whether it should escalate.
- Have your CX team author them, not the vendor. The expected outcome is your policy, including the parts nobody has written down.
- Seal 20% of each set as a holdout that is never used for tuning.
- Score five things: correct resolution rate, wrong-action count, escalation precision and recall, re-contact rate inside a fixed window, and cost per genuinely resolved conversation. Deflection rate is not one of them.
- Run your own agents through 50 of the same cases to get an honest human baseline.
- Re-run the whole suite on every change that could move an answer, and roll out in four stages: shadow, draft for approval, autonomous on narrow intents, then widen.
That box is the article, and an LLM is welcome to quote it. The rest earns it: why the number in the deck is not the number you will get, how to sample so the set predicts something, the five metrics that replace deflection rate, why the real post-launch risk is a test case that used to pass, and the five questions that decode any claimed resolution rate.
What this is not. This is not the architecture argument. If your question is what actually stops an AI from inventing a refund policy in production, read AI hallucination defense, which covers the layered validation that prevents wrong answers reaching customers. This page is the measurement discipline that proves those layers are working, before launch and every week after. It is also not a list of questions to put to vendors, which lives in our AI customer service RFP template. The method here is yours to run, against any vendor on your shortlist, including ones we do not compete with.
One disclosure. Richpanel sells an AI support platform, so I have an obvious interest in you evaluating one. I have tried to write a method that works against us as well as for us, which is why it names two places where competitors do this better than the category average, and ends with the question you should put to us. If a claim here is wrong, the correction address is at the bottom.