Product

AI AgentsAI SafetyHelpdeskSelf-Service PortalSocial Media AIReturns & Exchanges

Results

Customer Stories50% Performance GuaranteeTrust Center

More

PricingMigrationPartnersLog in
For Heads of CX and VPs of Support running 50 to 100+ agents

At 70 agents, your next 20 hires buy less capacity than you think.

Volume is climbing, headcount is capped, and both obvious answers disappoint. Hiring delivers a fraction of the capacity the plan promises once ramp, attrition backfill, and span of control are priced in. Buying a bot and pointing it at the whole queue underperforms in every category at once. What works at your scale is to re-underwrite the queue category by category against four treatments, then insist the automated band can take real actions in your systems. Here is the capacity math, the treatment map, the guardrail spec to hand vendors before the demo, and the scorecard that catches one gaming it.

By Amit RG, Founder, Richpanel Published 2026-08-04 ~14 min read
View as Markdown →
AR
Amit RG is the founder of Richpanel, the AI-native helpdesk serving 3,000+ brands. The capacity model below is a planning model with its assumptions stated, not measured data. The customer outcomes are published with written permission and are cited with their real scope. Competitor billing facts are stated as of August 2026 and are drawn from the vendors' own published pricing. On X: @realamitrg.
The short answer

How to scale support without hiring at 50 to 100+ agents

Support capacity is three numbers multiplied together: how many contacts arrive, what share of them need a human, and how many minutes a human spends on each. Hiring only moves the last one's denominator, and at your scale it moves it less than the plan assumes, because ramp time, attrition backfill, and one more team lead per eight to ten agents eat most of every net new head. The number that changes the shape is the middle one, the human contact rate. You lower it by sorting every ticket category into exactly one of four treatments, eliminate, self-serve, automate, or assist, and then holding the automated band to a hard standard: it has to take the action, not describe it. An AI that looks up the order, issues the refund inside a monetary cap, changes the delivery address before the fulfilment cutoff, and pauses the subscription removes a contact from your capacity model. An AI that paraphrases your help center moves that contact somewhere else and often creates a second one.

The math that decides the plan

Capacity is three numbers. Hiring only moves one of them.

Before any vendor conversation, plot your own ceiling. The identity is simple enough to run in a spreadsheet, and running it is what turns "we need more people" into a specific, defensible number.

Monthly human capacity = agents × handling hours per agent per month × (60 ÷ average handle time in minutes). And on the other side of the ledger, demand on humans = contacts arriving × the share that needs a human. Everything a CX leader can do sits in one of those four terms.

Work it with planning defaults, then replace every default with your own workforce-management data. A full-time agent is paid for about 173 hours a month (2,080 a year divided by twelve). Shrinkage, meaning holiday, sick time, training, coaching, meetings, and breaks, commonly runs 25% to 35%, leaving roughly 121 on-the-floor hours at a 30% assumption. Occupancy, the share of floor time actually spent handling contacts, typically runs 75% to 85%. At 80% that is about 97 handling hours per agent per month. Divide by handle time and you get the number most support plans never write down.

Average handle timeContacts per agent / monthA 70-agent team absorbs
8 minutes (simple email, chat)~730~51,000
10 minutes (mixed queue)~580~41,000
12 minutes (complex, multi-touch)~485~34,000

Model, not measurement. Assumes 173 paid hours per FTE per month, 30% shrinkage, 80% occupancy. Substitute your own three inputs before you use any of these numbers in a plan.

What the marginal hire actually delivers

The plan usually assumes a new agent contributes a full 730 contacts a month. Three things happen instead, and all three get worse as the floor gets bigger.

All three ranges are planning conventions, not measurements. Pull your own ramp curve from your last ten hires and your own attrition from HR before either number goes into a business case.

Which is why the four levers are not equal. This is the table to argue over with your CFO, because it reframes the question from "how many more agents" to "which term are we actually moving."

LeverHow you move itWhat it does to the ceiling
Contacts arrivingFix the upstream cause: shipping delays, unclear product pages, checkout errors, silent trackingRemoves demand outright. The cheapest capacity you will ever buy.
Share needing a humanSelf-serve for the simple cases, AI that takes real actions for the restThe only structural lever. Lowers demand without touching quality or headcount.
Minutes per contactCopilot drafting, better tooling, unified customer context, macrosRaises throughput per head. Does not change the shape of the curve.
Number of agentsHiringMost expensive, slowest, and partly consumed by backfill and ramp.
The framework

Four treatments. Every category gets exactly one.

The most common failure at 50 to 100+ agents is not picking the wrong vendor. It is applying one treatment to a queue that needs four, then blaming the AI rather than the assignment. Sort every category above roughly 3% of volume into one of these, and only one.

A fifth bucket exists and should be named explicitly so nobody automates into it by accident: human only. Complaints with legal or safety exposure, regulated claims, and anything where being wrong is expensive in a way a refund cannot fix.

The treatment map, applied

Volume shares below are the ranges we see across ecommerce and subscription queues. Yours will differ, which is the point of tagging your own week first.

CategoryPrimary treatmentWhat it takes, and how it fails
Order status and WISMO
20% to 40% of queue
Eliminate first, then self-serve, then automate the exceptionsProactive delay notifications and accurate tracking kill most of it. A portal lookup takes the next slice. AI handles what is left: the stuck, late, and lost parcels that need a carrier read plus a judgment call. Fails when you jump straight to automating it, because the contact still gets created and, on a per-resolution meter, still gets billed.
Returns and exchanges
10% to 20%
Self-serve portal, automate the policy exceptionsA portal with a policy engine and label generation handles the clean cases. Fails when the portal cannot express your real policy, so every edge case becomes a ticket anyway and you have paid for a portal that raised the average complexity of your queue.
Refunds and store credit
5% to 15%
Automate, under a monetary ceilingWrite access to the order or payment system, a per-action cap, a daily aggregate cap, and an audit log. Fails when the AI can only explain the refund policy, which means a human still does the work and you have automated the easy half of a two-step job.
Cancellations and subscription changes
10% to 25% on subscription brands
Automate, with a save flowWrite access to the billing or subscription system so pause, skip, swap, frequency, and address changes actually execute. Fails when the integration is read-only: the AI answers correctly and the human still clicks the button. This is also the one category that is a revenue lever, not just a cost one.
Delivery address edits
3% to 8%
Automate, time-boxedA write to the order before the fulfilment cutoff, and a hard stop after it. Fails when there is no cutoff logic, so the AI cheerfully promises a change on a parcel the warehouse already shipped.
Product, sizing, compatibility
10% to 20%
AutomateLive catalog and policy grounding, your brand voice, and an explicit behaviour when the answer is not in scope. Fails when the catalog is not connected and the model fills the gap confidently.
Damaged, defective, missing items
5% to 10%
Assist, then automate the clean casesEvidence capture first, then a bounded replacement or credit for the unambiguous cases. Fails when you automate the judgment cases and take the CSAT hit on your most emotional contacts.
Billing disputes and chargebacks
1% to 5%
Assist and escalateThe AI assembles the transaction history and evidence pack; a human decides. Fails when automated at all.
Complaints, legal, safety, regulated claims
1% to 3%
Human onlyRouting that recognises them on arrival, which is itself a good use of AI. Fails when the classifier is tuned for coverage instead of caution.
VIP, wholesale, B2B accounts
Varies
AssistFull account context in one pane for a named owner. Fails when treated as routine volume.

Two things usually surprise people the first time they run this. The eliminate and self-serve columns are larger than expected, and they are the cheapest capacity available. And the automate column is almost entirely gated on write-actions, not on how clever the model is.

The distinction that decides everything

An AI that answers moves work around. An AI that acts removes it.

This is the single test that predicts whether a deployment shows up in your capacity model or only in a dashboard. A contact that is contained but not completed does not leave the system. It comes back as a follow-up, which is two contacts where the plan counted zero, or it churns quietly, which is worse and invisible.

Run the test category by category. For each of your top ten categories by volume, write down the specific write-call that ends the conversation, then ask whether the AI can make it.

The contactAn answering AIAn acting AI
"Cancel my subscription"Links to the account page and explains the policyAsks the reason, offers a pause or a discount if the save flow allows it, then executes the change in the billing system and confirms
"This arrived broken"Describes the returns policyCaptures the photo, checks the order and eligibility, issues the replacement or credit within its cap, and sends the label
"Change my delivery address"Says addresses can be changed before dispatchChecks the fulfilment status against the cutoff, writes the new address if it is still open, and escalates if it is not
"Where is my order"Repeats the tracking link the customer already hasReads the carrier status, recognises a stalled scan, and either reships or explains the real delay with a date

The left column is a help-center search with better manners. It genuinely resolves informational contacts, and those are real, so do not dismiss the class. But in a mature ecommerce or subscription queue they are the minority, and a vendor whose demo lives entirely in that column is showing you the easy half.

CX leaders describe the resulting failure in almost identical words. One told us their team ended up having to "babysit it and fine-tune it by doing manual QA on AI tickets." That is what an answering bot costs you: the contact is not removed, and now a human reviews the bot as well. Names withheld; these are prospects, not published customers. The longer version of this distinction is in AI chatbot versus AI agent, and the cost-per-ticket cut for smaller teams is in the DTC playbook.

The control spec

Guardrails are an operating requirement, not a feature list.

The moment an AI can move money and change orders, it is an operator with system access, and it should be governed like one. Write these seven controls down as a document and send it to vendors before the first demo. It reorders the entire conversation from features to controls, and it is the fastest way to find out who has actually run this at scale.

01

Identity before any account action.

The AI authenticates the requester against the order or account before it reads personal data and before it changes anything.

This is where most deployments are weakest, because in a chat window the customer feels already identified and usually is not. Set the verification standard per action class: reading order status is not the same risk as changing a shipping address, which is not the same risk as refunding to a payment method.

The question to ask: What identity check runs before each action class, and what does the AI do when verification fails?
02

Monetary and permission ceilings, at three levels.

Per action, per conversation, and per day in aggregate. Above the ceiling the AI escalates rather than deciding.

A per-action cap alone is not enough at your volume. The failure you are protecting against is not one large refund, it is a policy misreading applied at scale for six hours before anyone notices. The daily aggregate cap is the circuit breaker, and most vendors do not ship it until an enterprise asks.

The question to ask: Can I set per-action, per-conversation, and daily aggregate limits myself, and who gets alerted when one trips?
03

An enumerated action inventory.

The AI can call only the tools you enabled, with only the arguments you allow. No free-form writes into production systems.

Deterministic tool execution is what separates an agent from an improviser. The model picks which enumerated action fits; it does not compose novel operations against your order system. It also makes the deployment reviewable, because a finite list of actions is something security and finance can sign off on.

The question to ask: Show me the full list of write-actions the AI can perform in my systems, and how I turn each one on or off.
04

Explicit unknown behaviour, and escalation on ambiguity.

The AI must be willing to stop. A confident answer outside its knowledge is worse than a handoff.

One CX leader at a safety-critical parts brand put the requirement better than any spec we have written: any AI they use has to be "willing to give up easily... if it's reached the bound of its legitimate knowledge." Another asked what every buying committee eventually asks: "What is the hallucination rate? I wouldn't believe that it had never made a mistake." The right answer is a mechanism, not a number: grounding in your content, a second model reviewing before send, and a defined escalation path. Full architecture in the hallucination defense piece.

The question to ask: Show me three conversations where the AI escalated instead of answering, and tell me what triggered each one.
05

Audit trail and rollback.

Every action logged with its inputs, its reasoning, and its result. Reversible wherever the underlying system allows a reversal.

Three audiences need this: the customer disputing what happened, finance reconciling refunds, and the auditor checking that access controls work. Ask for the export format early. "It is in the conversation view" is not an audit trail.

The question to ask: Can I export a full action log for a date range, and which actions can be reversed from inside the platform?
06

Data boundaries, residency, and retention.

What the AI can see, where that data sits, how long it is kept, and which sub-processors touch it.

At 50 to 500 employees this stops being a CX decision and becomes a security review. One buying committee compressed the whole list into a sentence: "where does PII sit... SOC 2, SOC 3 stuff... how sustainable are you guys?" Have the answers before the review rather than during it, and ask every vendor for the reports, not the badge.

The question to ask: Send me the current SOC 2 Type II report, the sub-processor list, and the retention policy, under NDA if needed.
07

Regression protection before every change, and copilot before autonomy.

Test cases authored by your CX team, run before any policy or prompt change ships. Autonomy released category by category on measured accuracy, never as a single switch.

Running the AI as a copilot first is a control, not a timidity phase. Agents approve drafts, you watch accuracy per category on real traffic, and you release autonomy where the numbers hold. The throughput gain arrives during that period anyway. Insist the test cases live on your side: if all you can see is a vendor dashboard score, you cannot audit a regression, and you will hear about one from a customer.

The question to ask: Who writes the test cases, can I add my own, and can I see the diff when a policy change ships?
What to measure

Seven numbers worth tracking, and four that will flatter you.

Agree the scorecard and the definitions in writing before the pilot starts. Half of the disappointing AI deployments we hear about were measured on a metric that rose while the queue did too.

MetricDefinition to insist onWhy it matters at your scale
Resolution rate by categoryClosed by AI with no human reply and no reopen inside a stated window, as a share of conversations in that categoryThe only metric that maps to your capacity model. By category, because the average hides which bands are actually working.
Reopen rate on AI-closed conversationsReopened or followed up within 7 daysThe honesty check on the number above. A high resolution rate with a rising reopen rate is a relabelling exercise.
Escalation rate, and escalation precisionShare handed to a human, plus the share of those that a human agrees needed a humanEscalating too little is a risk problem. Escalating too much is a cost problem. Only tracking both tells you which one you have.
CSAT, split three waysAI-resolved, human-resolved, and human-after-AIThe third split is where the damage hides. If conversations rescued from the AI score badly, the AI is buying capacity with goodwill.
Cost per genuinely resolved conversation(All AI platform and per-resolution fees + the fully-loaded human minutes spent on AI-touched conversations) ÷ conversations closed with no human reply and no reopenCatches double-metering and catches an AI that "resolves" while quietly generating human rework. Vendors rarely volunteer this one.
Capacity released, in agent hoursContacts removed from humans × your average handle timeThe number your WFM plan and your CFO consume. Divide by about 97 handling hours to express it as full-time equivalents.
Guardrail eventsPolicy violations caught before send, versus reached the customerThe ratio between those two is your real safety margin, and it should improve month over month.

The four that will flatter you

None of these are dishonest metrics in themselves. They become dishonest when they are the pilot success criteria. Put resolution rate with the reopen window stated at the top of the contract, and the rest can stay on the dashboard.

The billing architecture

At your volume, the meter matters more than the rate card.

A price difference that looks like rounding at 2,000 conversations a month is a hiring-plan-sized number at 100,000. Two questions decide it: are you billed once or twice for one AI-resolved conversation, and what happens when the allowance runs out.

Take 50,000 AI-resolved conversations a month, a realistic band for a 70-agent operation once the automate treatments are live. At a $1.00 per-resolution meter that is $50,000 a month, or $600,000 a year, for the resolutions alone. At roughly $0.20 it is $10,000 a month, or $120,000. The $480,000 gap is the difference between the programme funding itself and the programme being the reason the software line went up.

Published and publicly reported list pricing, checked August 2026. Contact-us vendors are shown at their publicly reported ranges, which is not the same thing as a quote.

PlatformHow AI is billedThe thing to check at scale
Gorgias$0.90 per automated interaction in-allowance, $1.50 beyond, on top of ticket-based helpdesk pricingTheir own pricing tooltip states that an automated interaction resolving without a human within 72 hours also counts as one helpdesk ticket. On their published $1,227 card (5,000 tickets plus 530 interactions: helpdesk $750, AI Agent $477) that is roughly $1.05 per AI-resolved ticket inside the allowance, and about $1.86 past it once the $0.36 ticket and $1.50 interaction overages apply.
ZendeskAutonomous AI is bundled across plans, with per-resolution charges on top of seatsThe add-on stack is the enterprise story: roughly $115 a seat, plus about $50 for Copilot, plus about $50 for QA and workforce management, lands near $215 a seat before any AI resolution charges. Real-world resolution on their native AI lands well below the marketed range.
Intercom Fin$0.99 per resolution on top of $29 to $132 seatsSalesforce signed an agreement to acquire Fin in June 2026, pending close. Not a reason to rule it out, but a roadmap question worth asking directly.
Freshdesk Freddy$49 per 100 sessions, plus about $29 a seat for copilotBilled per session rather than per resolution, and the AI stops responding when sessions are exhausted. At enterprise volume that is a capacity cliff, not a billing detail.
Help Scout$0.75 per AI resolution on top of seatsSame double-count question as any per-ticket helpdesk with a resolution meter bolted on.
Salesforce AgentforceBoth a $2 per conversation model and Flex Credits at roughly $0.10 per action are liveRequires Service Cloud and Data Cloud, so the real number is a platform TCO question, not an AI price question.
DecagonOutcome-based, contact-us pricing; publicly reported in the $95K to $590K+ rangeNo double meter, since it is not a helpdesk. You are still paying for the helpdesk it runs on, plus separate QA and knowledge tooling.
SierraOutcome-based, roughly a $150K floorTheir own published case studies put deployment in the 4 to 10 week range. Same second-bill point as above.
LorikeetPublic pay-per-resolution, roughly $1,500 to $4,000 a monthStrongest in complex and regulated queues. Again an agent layer, not a helpdesk.
RichpanelAbout $0.20 per AI-resolved conversation on a $200 a month minimum, plus $99 per human seat on an annual commitment ($119 billed monthly, $129 on demand)Metered once. The seat line is per human, so it does not grow with volume; the conversation line is the one that scales, and you choose the model and token budget behind it.

Marginal savings and blended savings are different numbers

Every multiple above is marginal: one AI-resolved conversation against one human-resolved conversation. About $0.20 against a roughly $3 fully-loaded human conversation is about 15x, and against a $0.90 to $2 competitor AI meter it is around 5x. Your blended total cost of support falls far less, because you keep most of your team, you keep the platform fees, and the categories you did not automate cost exactly what they cost today. In a model that automates half the volume and retains most of the floor, blended savings land in the tens of percent, not multiples. Take the blended number to your CFO on your real baseline. Use the marginal number only as the per-conversation proof inside it. Any vendor presenting a marginal multiple as a total-savings figure is doing arithmetic you should check.

The published deployments

Three deployments, cited with their real scope.

These are named customers published with written permission. They are placed against the treatment they actually ran, because the most common way proof gets misused in this category is by presenting a self-service result as an autonomous-AI result.

Aeons · the automate band

AI sends 63% of every customer message, at higher CSAT than the team.

Across the measured window the AI sent 63% of all customer messages and scored 4.39 out of 5 against 4.33 team-wide. The split underneath is the part worth copying: routine conversations close end to end with no human touch, and on the higher-risk cases the AI leads and a human reviews before sending. Capacity returned was about five full-time agents. That is the two-layer pattern the framework predicts, an autonomous band and a reviewed band split by risk rather than by ticket count. Full detail in the wellness case study.

Jones Road · the eliminate, self-serve, and assist bands

Tickets per order roughly halved, and the team went from 18 agents to 10.

Measured in their own analytics after migrating in September 2024: tickets per order fell from 22% to 25% down to 11% to 14%, the self-service portal resolves 37.5% of website and chat sessions (2,451 of 6,521), and CSAT sits at 4.2 out of 5. The team went from 18 agents to 10 with order volume flat or higher. Being precise matters here: Jones Road runs the helpdesk, the self-service portal, and an agent copilot. None of that is autonomous AI resolution, which is exactly why we cite it in this article. Most enterprise teams have more capacity sitting in the eliminate, self-serve, and assist bands than they expect, and it is available without ever handing a customer conversation to an autonomous agent.

Ridge · the cost outcome

Cost per ticket down about 70%, and CSAT up from 88% to 96%.

Roughly $500K in annual savings, with quality moving up over the same period. The pairing is the point. The default assumption in a headcount-constrained support organisation is that cheaper means worse, and it does not hold here because the volume moved off humans is the repetitive band, not the band where human judgment was creating the quality.

The honest note about scale

Our published references run support teams smaller than yours.

The deployments we can name with written permission operate support teams below the 70-agent mark. The capacity mechanism does not change with size, but the risks that decide a rollout at 50 to 100+ agents do: integration depth into systems that are rarely off-the-shelf at your size, a security review with a real committee behind it, and change management across a large floor. So ask us, and ask every vendor on your shortlist, for a reference at your headcount and in your vertical, before the technical evaluation rather than after. If a vendor cannot produce one, that is information, and it should change how you structure the pilot rather than end the conversation.

Where we fit

What Richpanel is, and where it is the wrong call.

Stated once, so the rest of this page can be read on its own merits.

Richpanel is an AI-native helpdesk: the AI agents that resolve the repeat work plus the helpdesk your humans work in, on a single meter. A CX Manager AI reads your past tickets, help center, and public reviews, then drafts the policies, escalation rules, and quality rubric, so setup is a configuration exercise rather than a consulting engagement. Your CX team owns the test cases, and a quality agent reviews closed conversations and feeds gaps back into policy. Time to value runs in three stages: a 30-minute proof of concept built live on your own data during the demo, a two-week pilot, then a four-week deployment. The de-risking is contractual, a 50% resolution guarantee in 30 days with your money back if we miss. We hold a SOC 2 Type II report with an unqualified opinion and zero exceptions across roughly 90 controls for the May to October 2025 period, plus independent HIPAA and GDPR assessments. Around 3,000 brands run on the platform. Pricing is in the pricing breakdown.

Where it is the wrong call

The honest caveats

What this approach cannot do.

What to do next

Five moves, in order, before you talk to a vendor.

Every one of these is doable in 30 days with the data you already have, and doing them first changes what a vendor can sell you.

01

Tag one full week of contacts by category.

Record volume and average handle time per category. This one artifact is the input to everything else on this page, and almost no team has it in usable form when the vendor calls start.

02

Compute your capacity identity from your own numbers.

Handling hours per agent per month, contacts per agent, utilisation, and the real cost of a net new head after ramp and backfill. Now you have a ceiling and a marginal cost, both defensible in a budget conversation.

03

Assign one treatment to every category above 3% of volume.

Eliminate, self-serve, automate, assist, or human only. For every category marked automate, name the write-call that ends the conversation and the system it lives in. That list is your integration requirement, and it is worth more in a vendor conversation than any feature checklist.

04

Write the seven-control guardrail spec and send it out first.

Before the demos. Vendors respond to a control document very differently than to an open discovery call, and those differences are the most useful signal you will get all quarter. A fuller structure sits in our vendor RFP template.

05

Fix the scorecard and the reopen window in writing before the pilot.

Resolution rate by category with the reopen window stated, reopen rate, escalation precision, three-way CSAT, and cost per genuinely resolved conversation. Agreeing these at week zero stops the goalposts moving at week four.

Seven questions that separate vendors quickly

Frequently asked

The questions large support teams ask first.

How do you scale support without hiring when you already have 50 to 100 agents?

You lower the share of contacts that need a human rather than raising the number of humans. Sort every ticket category above 3% of volume into one of four treatments: eliminate the cause, move it to self-serve, automate it end to end, or assist the human who handles it. Then hold the automated band to a hard standard, that it can take the write-action that closes the conversation, not just describe it. Hiring is the weakest of the four levers at your scale because ramp time, attrition backfill of 30% to 40% a year, and one more team lead per eight to ten agents consume most of every net new head before it delivers.

How many tickets can one support agent actually handle in a month?

Work it from handling hours rather than from a headline number. A full-time agent is paid about 173 hours a month. Take off shrinkage of 25% to 35% for holiday, sick time, training, coaching, meetings, and breaks, then apply occupancy of 75% to 85% for the share of floor time actually spent on contacts. That lands near 97 handling hours. At an 8-minute average handle time that is roughly 730 contacts a month; at 12 minutes it is about 485. A 70-agent team therefore absorbs somewhere between roughly 34,000 and 51,000 contacts a month. Substitute your own three inputs before using any of this in a plan.

What share of tickets can AI realistically resolve at enterprise scale?

It is set by two things: what share of your queue is repeat work, and how many of your resolving actions the AI can actually execute. A mature deployment on a repeat-heavy ecommerce or subscription queue lands in the 50% to 80% band, and 50% in 30 days is the number we guarantee contractually with money back if we miss. But that ceiling is capped by integration depth, not by the model. If the calls that close your tickets live in a custom order system the AI cannot reach, those categories stay with humans regardless of how good the AI is. Ask any vendor for resolution rate by category with a reopen window stated, because the average hides which bands are working.

Should we run AI as a copilot first, or go straight to autonomous?

Copilot first, per category, and treat it as a control rather than a timidity phase. The AI drafts, your agents approve, you watch accuracy on your real traffic by category, and you release autonomy where the numbers hold. You get the throughput gain during that period anyway, and you get an evidence base instead of a leap. The failure this prevents is the one everybody has heard about: a blind cutover, one bad week, and an AI programme that gets switched off permanently for a reason that was a configuration problem.

How do we stop an AI agent from issuing a refund it should not?

With controls, not with confidence. Identity verification before any account action. Monetary ceilings at three levels: per action, per conversation, and per day in aggregate, because the risk at your volume is not one large refund but a policy misreading applied at scale for hours. An enumerated action inventory, so the AI can only call tools you enabled with arguments you allow, rather than composing operations against production. Explicit escalation when it is out of scope. A full audit log with reversal where the underlying system allows it. And test cases your team owns, run before any policy change ships. Ask a vendor to demonstrate all six, and ask to see three real conversations where the AI stopped and escalated.

What should we measure in an AI support pilot?

Seven numbers: resolution rate by category with the reopen window stated, reopen rate on AI-closed conversations within 7 days, escalation rate paired with escalation precision, CSAT split three ways (AI-resolved, human-resolved, and human-after-AI), cost per genuinely resolved conversation including the human minutes spent on AI-touched conversations, capacity released in agent hours, and guardrail events caught before send versus reached the customer. Keep containment rate, deflection rate, AI-touched share, and first response time off the success criteria. All four can improve while your queue grows.

Bring your top ten categories. We will show you which ones resolve.

On a 30-minute call we build a proof of concept live on your own data and run it against the tickets you find hardest, with the write-actions connected. Then a two-week pilot, then full deployment in four. If the AI is not resolving half your volume in 30 days, you get your money back.

Book a 30-min demo →