The short answer
At high volume, four disciplines decide whether AI support works. None of them is model selection, and each has a measurable standard and a full guide below.
- Capacity. Know your own arithmetic before a vendor shows you theirs. Hiring is the weakest lever you have.
- Evaluation. The only resolution rate that predicts anything was measured on your tickets. Build the test set first.
- QA. AI mistakes are correlated where human mistakes are independent. A 2% sample that was adequate for people is blind for machines.
- Governance. Prompts are instructions, not permissions. Cap what the AI can do in your systems, not what it is told.
Each figure is worked in full, with assumptions and sources, in the guide its section links below.