---
title: "How to Scale Support Without Hiring at 50 to 100+ Agents"
description: "A capacity playbook for Heads of CX running 50 to 100+ agents: the capacity model, the eliminate / self-serve / automate / assist treatment map applied to real ticket categories, the seven-control guardrail spec, the billing architectures compared at enterprise volume, and the scorecard that catches a vendor gaming it."
url: https://www.richpanel.com/learn/scale-support-without-hiring-enterprise
datePublished: 2026-08-04
dateModified: 2026-08-04
author: "Amit RG"
source: richpanel.com
---

# At 70 agents, your next 20 hires buy *less capacity than you think.*

Volume is climbing, headcount is capped, and both obvious answers disappoint. Hiring delivers a fraction of the capacity the plan promises once ramp, attrition backfill, and span of control are priced in. Buying a bot and pointing it at the whole queue underperforms in every category at once. What works at your scale is to re-underwrite the queue category by category against four treatments, then insist the automated band can take real actions in your systems. Here is the capacity math, the treatment map, the guardrail spec to hand vendors before the demo, and the scorecard that catches one gaming it.

> **Amit RG** is the founder of Richpanel, the AI-native helpdesk serving 3,000+ brands. The capacity model below is a planning model with its assumptions stated, not measured data. The customer outcomes are published with written permission and are cited with their real scope. Competitor billing facts are stated as of August 2026 and are drawn from the vendors' own published pricing. On X: [@realamitrg](https://x.com/realamitrg).

#### How to scale support without hiring at 50 to 100+ agents

Support capacity is three numbers multiplied together: how many contacts arrive, what share of them need a human, and how many minutes a human spends on each. Hiring only moves the last one's denominator, and at your scale it moves it less than the plan assumes, because ramp time, attrition backfill, and one more team lead per eight to ten agents eat most of every net new head. The number that changes the shape is the middle one, the human contact rate. You lower it by sorting every ticket category into exactly one of four treatments, **eliminate, self-serve, automate, or assist**, and then holding the automated band to a hard standard: it has to take the action, not describe it. An AI that looks up the order, issues the refund inside a monetary cap, changes the delivery address before the fulfilment cutoff, and pauses the subscription removes a contact from your capacity model. An AI that paraphrases your help center moves that contact somewhere else and often creates a second one.

## Capacity is three numbers. Hiring only moves *one of them.*

Before any vendor conversation, plot your own ceiling. The identity is simple enough to run in a spreadsheet, and running it is what turns "we need more people" into a specific, defensible number.

**Monthly human capacity = agents × handling hours per agent per month × (60 &divide; average handle time in minutes).** And on the other side of the ledger, **demand on humans = contacts arriving × the share that needs a human.** Everything a CX leader can do sits in one of those four terms.

Work it with planning defaults, then replace every default with your own workforce-management data. A full-time agent is paid for about 173 hours a month (2,080 a year divided by twelve). Shrinkage, meaning holiday, sick time, training, coaching, meetings, and breaks, commonly runs 25% to 35%, leaving roughly 121 on-the-floor hours at a 30% assumption. Occupancy, the share of floor time actually spent handling contacts, typically runs 75% to 85%. At 80% that is **about 97 handling hours per agent per month.** Divide by handle time and you get the number most support plans never write down.

| Average handle time | Contacts per agent / month | A 70-agent team absorbs |
| --- | --- | --- |
| 8 minutes (simple email, chat) | ~730 | ~51,000 |
| 10 minutes (mixed queue) | ~580 | ~41,000 |
| 12 minutes (complex, multi-touch) | ~485 | ~34,000 |

Model, not measurement. Assumes 173 paid hours per FTE per month, 30% shrinkage, 80% occupancy. Substitute your own three inputs before you use any of these numbers in a plan.

### What the marginal hire actually delivers

The plan usually assumes a new agent contributes a full 730 contacts a month. Three things happen instead, and all three get worse as the floor gets bigger.

- **Ramp.** Four to eight weeks to full productivity is a common planning assumption for a mixed ecommerce queue. Across a first quarter that is roughly half a head of delivered capacity, not one.
- **Attrition backfill.** Frontline support attrition frequently runs 30% to 40% a year. On a 70-agent floor that is 21 to 28 people you hire annually just to stay level. A plan to reach 90 agents is really a plan to hire around 45 people in a year, and to onboard all of them.
- **Span of control.** One team lead per eight to ten agents, plus quality-assurance capacity that in most centres manually samples a low single-digit percentage of conversations. Twenty more agents means two or three more leads and more QA coverage. Headcount cost grows faster than headcount.

All three ranges are planning conventions, not measurements. Pull your own ramp curve from your last ten hires and your own attrition from HR before either number goes into a business case.

Which is why the four levers are not equal. This is the table to argue over with your CFO, because it reframes the question from "how many more agents" to "which term are we actually moving."

| Lever | How you move it | What it does to the ceiling |
| --- | --- | --- |
| **Contacts arriving** | Fix the upstream cause: shipping delays, unclear product pages, checkout errors, silent tracking | Removes demand outright. The cheapest capacity you will ever buy. |
| **Share needing a human** | Self-serve for the simple cases, AI that takes real actions for the rest | **The only structural lever.** Lowers demand without touching quality or headcount. |
| **Minutes per contact** | Copilot drafting, better tooling, unified customer context, macros | Raises throughput per head. Does not change the shape of the curve. |
| **Number of agents** | Hiring | Most expensive, slowest, and partly consumed by backfill and ramp. |

## Four treatments. Every category gets *exactly one.*

The most common failure at 50 to 100+ agents is not picking the wrong vendor. It is applying one treatment to a queue that needs four, then blaming the AI rather than the assignment. Sort every category above roughly 3% of volume into one of these, and only one.

- **Eliminate.** The contact is never created because the cause is fixed: proactive delay notifications, accurate tracking, clearer sizing information, a checkout that does not fail. This work sits with product and operations, which is exactly why it never happens unless a CX leader pushes it with data.
- **Self-serve.** The customer completes it themselves in a portal: track the order, start a return, swap a size, edit the delivery address, pause a subscription. No conversation is created at all.
- **Automate.** A conversation is created and the AI closes it end to end, including the write-action in your systems. Strictest requirements, largest payoff.
- **Assist.** A human handles it faster, with AI drafting the reply and assembling the account context. Throughput, not deflection.

A fifth bucket exists and should be named explicitly so nobody automates into it by accident: **human only.** Complaints with legal or safety exposure, regulated claims, and anything where being wrong is expensive in a way a refund cannot fix.

### The treatment map, applied

Volume shares below are the ranges we see across ecommerce and subscription queues. Yours will differ, which is the point of tagging your own week first.

| Category | Primary treatment | What it takes, and how it fails |
| --- | --- | --- |
| **Order status and WISMO** 20% to 40% of queue | Eliminate first, then self-serve, then automate the exceptions | Proactive delay notifications and accurate tracking kill most of it. A portal lookup takes the next slice. AI handles what is left: the stuck, late, and lost parcels that need a carrier read plus a judgment call. **Fails when** you jump straight to automating it, because the contact still gets created and, on a per-resolution meter, still gets billed. |
| **Returns and exchanges** 10% to 20% | Self-serve portal, automate the policy exceptions | A portal with a policy engine and label generation handles the clean cases. **Fails when** the portal cannot express your real policy, so every edge case becomes a ticket anyway and you have paid for a portal that raised the average complexity of your queue. |
| **Refunds and store credit** 5% to 15% | Automate, under a monetary ceiling | Write access to the order or payment system, a per-action cap, a daily aggregate cap, and an audit log. **Fails when** the AI can only explain the refund policy, which means a human still does the work and you have automated the easy half of a two-step job. |
| **Cancellations and subscription changes** 10% to 25% on subscription brands | Automate, with a save flow | Write access to the billing or subscription system so pause, skip, swap, frequency, and address changes actually execute. **Fails when** the integration is read-only: the AI answers correctly and the human still clicks the button. This is also the one category that is a revenue lever, not just a cost one. |
| **Delivery address edits** 3% to 8% | Automate, time-boxed | A write to the order before the fulfilment cutoff, and a hard stop after it. **Fails when** there is no cutoff logic, so the AI cheerfully promises a change on a parcel the warehouse already shipped. |
| **Product, sizing, compatibility** 10% to 20% | Automate | Live catalog and policy grounding, your brand voice, and an explicit behaviour when the answer is not in scope. **Fails when** the catalog is not connected and the model fills the gap confidently. |
| **Damaged, defective, missing items** 5% to 10% | Assist, then automate the clean cases | Evidence capture first, then a bounded replacement or credit for the unambiguous cases. **Fails when** you automate the judgment cases and take the CSAT hit on your most emotional contacts. |
| **Billing disputes and chargebacks** 1% to 5% | Assist and escalate | The AI assembles the transaction history and evidence pack; a human decides. **Fails when** automated at all. |
| **Complaints, legal, safety, regulated claims** 1% to 3% | Human only | Routing that recognises them on arrival, which is itself a good use of AI. **Fails when** the classifier is tuned for coverage instead of caution. |
| **VIP, wholesale, B2B accounts** Varies | Assist | Full account context in one pane for a named owner. **Fails when** treated as routine volume. |

Two things usually surprise people the first time they run this. The eliminate and self-serve columns are larger than expected, and they are the cheapest capacity available. And the automate column is almost entirely gated on write-actions, not on how clever the model is.

## An AI that answers moves work around. An AI that *acts* removes it.

This is the single test that predicts whether a deployment shows up in your capacity model or only in a dashboard. A contact that is contained but not completed does not leave the system. It comes back as a follow-up, which is two contacts where the plan counted zero, or it churns quietly, which is worse and invisible.

Run the test category by category. For each of your top ten categories by volume, write down the specific write-call that ends the conversation, then ask whether the AI can make it.

| The contact | An answering AI | An acting AI |
| --- | --- | --- |
| "Cancel my subscription" | Links to the account page and explains the policy | Asks the reason, offers a pause or a discount if the save flow allows it, then executes the change in the billing system and confirms |
| "This arrived broken" | Describes the returns policy | Captures the photo, checks the order and eligibility, issues the replacement or credit within its cap, and sends the label |
| "Change my delivery address" | Says addresses can be changed before dispatch | Checks the fulfilment status against the cutoff, writes the new address if it is still open, and escalates if it is not |
| "Where is my order" | Repeats the tracking link the customer already has | Reads the carrier status, recognises a stalled scan, and either reships or explains the real delay with a date |

The left column is a help-center search with better manners. It genuinely resolves informational contacts, and those are real, so do not dismiss the class. But in a mature ecommerce or subscription queue they are the minority, and a vendor whose demo lives entirely in that column is showing you the easy half.

CX leaders describe the resulting failure in almost identical words. One told us their team ended up having to "babysit it and fine-tune it by doing manual QA on AI tickets." That is what an answering bot costs you: the contact is not removed, and now a human reviews the bot as well. Names withheld; these are prospects, not published customers. The longer version of this distinction is in [AI chatbot versus AI agent](https://www.richpanel.com/learn/ai-chatbot-vs-ai-agent), and the cost-per-ticket cut for smaller teams is in [the DTC playbook](https://www.richpanel.com/learn/scale-support-without-hiring-dtc).

## Guardrails are an operating requirement, *not a feature list.*

The moment an AI can move money and change orders, it is an operator with system access, and it should be governed like one. Write these seven controls down as a document and send it to vendors before the first demo. It reorders the entire conversation from features to controls, and it is the fastest way to find out who has actually run this at scale.

### 01. Identity before any account action.

The AI authenticates the requester against the order or account before it reads personal data and before it changes anything.

This is where most deployments are weakest, because in a chat window the customer feels already identified and usually is not. Set the verification standard per action class: reading order status is not the same risk as changing a shipping address, which is not the same risk as refunding to a payment method.

**The question to ask:** What identity check runs before each action class, and what does the AI do when verification fails?

### 02. Monetary and permission ceilings, at three levels.

Per action, per conversation, and per day in aggregate. Above the ceiling the AI escalates rather than deciding.

A per-action cap alone is not enough at your volume. The failure you are protecting against is not one large refund, it is a policy misreading applied at scale for six hours before anyone notices. The daily aggregate cap is the circuit breaker, and most vendors do not ship it until an enterprise asks.

**The question to ask:** Can I set per-action, per-conversation, and daily aggregate limits myself, and who gets alerted when one trips?

### 03. An enumerated action inventory.

The AI can call only the tools you enabled, with only the arguments you allow. No free-form writes into production systems.

Deterministic tool execution is what separates an agent from an improviser. The model picks which enumerated action fits; it does not compose novel operations against your order system. It also makes the deployment reviewable, because a finite list of actions is something security and finance can sign off on.

**The question to ask:** Show me the full list of write-actions the AI can perform in my systems, and how I turn each one on or off.

### 04. Explicit unknown behaviour, and escalation on ambiguity.

The AI must be willing to stop. A confident answer outside its knowledge is worse than a handoff.

One CX leader at a safety-critical parts brand put the requirement better than any spec we have written: any AI they use has to be "willing to give up easily... if it's reached the bound of its legitimate knowledge." Another asked what every buying committee eventually asks: "What is the hallucination rate? I wouldn't believe that it had never made a mistake." The right answer is a mechanism, not a number: grounding in your content, a second model reviewing before send, and a defined escalation path. Full architecture in [the hallucination defense piece](https://www.richpanel.com/learn/ai-hallucination-defense).

**The question to ask:** Show me three conversations where the AI escalated instead of answering, and tell me what triggered each one.

### 05. Audit trail and rollback.

Every action logged with its inputs, its reasoning, and its result. Reversible wherever the underlying system allows a reversal.

Three audiences need this: the customer disputing what happened, finance reconciling refunds, and the auditor checking that access controls work. Ask for the export format early. "It is in the conversation view" is not an audit trail.

**The question to ask:** Can I export a full action log for a date range, and which actions can be reversed from inside the platform?

### 06. Data boundaries, residency, and retention.

What the AI can see, where that data sits, how long it is kept, and which sub-processors touch it.

At 50 to 500 employees this stops being a CX decision and becomes a security review. One buying committee compressed the whole list into a sentence: "where does PII sit... SOC 2, SOC 3 stuff... how sustainable are you guys?" Have the answers before the review rather than during it, and ask every vendor for the reports, not the badge.

**The question to ask:** Send me the current SOC 2 Type II report, the sub-processor list, and the retention policy, under NDA if needed.

### 07. Regression protection before every change, and copilot before autonomy.

Test cases authored by your CX team, run before any policy or prompt change ships. Autonomy released category by category on measured accuracy, never as a single switch.

Running the AI as a copilot first is a control, not a timidity phase. Agents approve drafts, you watch accuracy per category on real traffic, and you release autonomy where the numbers hold. The throughput gain arrives during that period anyway. Insist the test cases live on your side: if all you can see is a vendor dashboard score, you cannot audit a regression, and you will hear about one from a customer.

**The question to ask:** Who writes the test cases, can I add my own, and can I see the diff when a policy change ships?

## Seven numbers worth tracking, and four that *will flatter you.*

Agree the scorecard and the definitions in writing before the pilot starts. Half of the disappointing AI deployments we hear about were measured on a metric that rose while the queue did too.

| Metric | Definition to insist on | Why it matters at your scale |
| --- | --- | --- |
| **Resolution rate by category** | Closed by AI with no human reply and no reopen inside a stated window, as a share of conversations in that category | The only metric that maps to your capacity model. By category, because the average hides which bands are actually working. |
| **Reopen rate on AI-closed conversations** | Reopened or followed up within 7 days | The honesty check on the number above. A high resolution rate with a rising reopen rate is a relabelling exercise. |
| **Escalation rate, and escalation precision** | Share handed to a human, plus the share of those that a human agrees needed a human | Escalating too little is a risk problem. Escalating too much is a cost problem. Only tracking both tells you which one you have. |
| **CSAT, split three ways** | AI-resolved, human-resolved, and human-after-AI | The third split is where the damage hides. If conversations rescued from the AI score badly, the AI is buying capacity with goodwill. |
| **Cost per genuinely resolved conversation** | (All AI platform and per-resolution fees + the fully-loaded human minutes spent on AI-touched conversations) &divide; conversations closed with no human reply and no reopen | Catches double-metering and catches an AI that "resolves" while quietly generating human rework. Vendors rarely volunteer this one. |
| **Capacity released, in agent hours** | Contacts removed from humans × your average handle time | The number your WFM plan and your CFO consume. Divide by about 97 handling hours to express it as full-time equivalents. |
| **Guardrail events** | Policy violations caught before send, versus reached the customer | The ratio between those two is your real safety margin, and it should improve month over month. |

### The four that will flatter you

- **Containment rate.** Counts a conversation as handled whether or not the customer's problem ended. Can rise while your queue grows.
- **Deflection rate.** Same problem, usually measured on help-center clicks, which is even further from an outcome.
- **"AI-touched" share.** A draft the human rewrote entirely still counts. Useful for a board slide, useless for a plan.
- **First response time.** An automated acknowledgement takes this to seconds without moving a single ticket out of the queue.

None of these are dishonest metrics in themselves. They become dishonest when they are the pilot success criteria. Put resolution rate with the reopen window stated at the top of the contract, and the rest can stay on the dashboard.

## At your volume, the meter matters *more than the rate card.*

A price difference that looks like rounding at 2,000 conversations a month is a hiring-plan-sized number at 100,000. Two questions decide it: are you billed once or twice for one AI-resolved conversation, and what happens when the allowance runs out.

Take 50,000 AI-resolved conversations a month, a realistic band for a 70-agent operation once the automate treatments are live. At a $1.00 per-resolution meter that is $50,000 a month, or $600,000 a year, for the resolutions alone. At roughly $0.20 it is $10,000 a month, or $120,000. The $480,000 gap is the difference between the programme funding itself and the programme being the reason the software line went up.

Published and publicly reported list pricing, checked August 2026. Contact-us vendors are shown at their publicly reported ranges, which is not the same thing as a quote.

| Platform | How AI is billed | The thing to check at scale |
| --- | --- | --- |
| **Gorgias** | $0.90 per automated interaction in-allowance, $1.50 beyond, on top of ticket-based helpdesk pricing | Their own pricing tooltip states that an automated interaction resolving without a human within 72 hours **also counts as one helpdesk ticket.** On their published $1,227 card (5,000 tickets plus 530 interactions: helpdesk $750, AI Agent $477) that is roughly $1.05 per AI-resolved ticket inside the allowance, and about $1.86 past it once the $0.36 ticket and $1.50 interaction overages apply. |
| **Zendesk** | Autonomous AI is bundled across plans, with per-resolution charges on top of seats | The add-on stack is the enterprise story: roughly $115 a seat, plus about $50 for Copilot, plus about $50 for QA and workforce management, lands near $215 a seat before any AI resolution charges. Real-world resolution on their native AI lands well below the marketed range. |
| **Intercom Fin** | $0.99 per resolution on top of $29 to $132 seats | Salesforce signed an agreement to acquire Fin in June 2026, pending close. Not a reason to rule it out, but a roadmap question worth asking directly. |
| **Freshdesk Freddy** | $49 per 100 sessions, plus about $29 a seat for copilot | Billed per session rather than per resolution, and the AI stops responding when sessions are exhausted. At enterprise volume that is a capacity cliff, not a billing detail. |
| **Help Scout** | $0.75 per AI resolution on top of seats | Same double-count question as any per-ticket helpdesk with a resolution meter bolted on. |
| **Salesforce Agentforce** | Both a $2 per conversation model and Flex Credits at roughly $0.10 per action are live | Requires Service Cloud and Data Cloud, so the real number is a platform TCO question, not an AI price question. |
| **Decagon** | Outcome-based, contact-us pricing; publicly reported in the $95K to $590K+ range | No double meter, since it is not a helpdesk. You are still paying for the helpdesk it runs on, plus separate QA and knowledge tooling. |
| **Sierra** | Outcome-based, roughly a $150K floor | Their own published case studies put deployment in the 4 to 10 week range. Same second-bill point as above. |
| **Lorikeet** | Public pay-per-resolution, roughly $1,500 to $4,000 a month | Strongest in complex and regulated queues. Again an agent layer, not a helpdesk. |
| **Richpanel** | About $0.20 per AI-resolved conversation on a $200 a month minimum, plus $99 per human seat on an annual commitment ($119 billed monthly, $129 on demand) | Metered once. The seat line is per human, so it does not grow with volume; the conversation line is the one that scales, and you choose the model and token budget behind it. |

#### Marginal savings and blended savings are different numbers

Every multiple above is **marginal**: one AI-resolved conversation against one human-resolved conversation. About $0.20 against a roughly $3 fully-loaded human conversation is about 15x, and against a $0.90 to $2 competitor AI meter it is around 5x. Your **blended** total cost of support falls far less, because you keep most of your team, you keep the platform fees, and the categories you did not automate cost exactly what they cost today. In a model that automates half the volume and retains most of the floor, blended savings land in the tens of percent, not multiples. Take the blended number to your CFO on your real baseline. Use the marginal number only as the per-conversation proof inside it. Any vendor presenting a marginal multiple as a total-savings figure is doing arithmetic you should check.

## Three deployments, cited with their *real scope.*

These are named customers published with written permission. They are placed against the treatment they actually ran, because the most common way proof gets misused in this category is by presenting a self-service result as an autonomous-AI result.

Aeons · the automate band

#### AI sends 63% of every customer message, at higher CSAT than the team.

Across the measured window the AI sent 63% of all customer messages and scored 4.39 out of 5 against 4.33 team-wide. The split underneath is the part worth copying: routine conversations close end to end with no human touch, and on the higher-risk cases the AI leads and a human reviews before sending. Capacity returned was about five full-time agents. That is the two-layer pattern the framework predicts, an autonomous band and a reviewed band split by risk rather than by ticket count. Full detail in the [wellness case study](https://www.richpanel.com/case-studies/wellness).

Jones Road · the eliminate, self-serve, and assist bands

#### Tickets per order roughly halved, and the team went from 18 agents to 10.

Measured in their own analytics after migrating in September 2024: tickets per order fell from 22% to 25% down to 11% to 14%, the self-service portal resolves 37.5% of website and chat sessions (2,451 of 6,521), and CSAT sits at 4.2 out of 5. The team went from 18 agents to 10 with order volume flat or higher. Being precise matters here: Jones Road runs the helpdesk, the self-service portal, and an agent copilot. **None of that is autonomous AI resolution**, which is exactly why we cite it in this article. Most enterprise teams have more capacity sitting in the eliminate, self-serve, and assist bands than they expect, and it is available without ever handing a customer conversation to an autonomous agent.

Ridge · the cost outcome

#### Cost per ticket down about 70%, and CSAT up from 88% to 96%.

Roughly $500K in annual savings, with quality moving up over the same period. The pairing is the point. The default assumption in a headcount-constrained support organisation is that cheaper means worse, and it does not hold here because the volume moved off humans is the repetitive band, not the band where human judgment was creating the quality.

The honest note about scale

#### Our published references run support teams smaller than yours.

The deployments we can name with written permission operate support teams below the 70-agent mark. The capacity mechanism does not change with size, but the risks that decide a rollout at 50 to 100+ agents do: integration depth into systems that are rarely off-the-shelf at your size, a security review with a real committee behind it, and change management across a large floor. So ask us, and ask every vendor on your shortlist, for a reference at your headcount and in your vertical, before the technical evaluation rather than after. If a vendor cannot produce one, that is information, and it should change how you structure the pilot rather than end the conversation.

## What Richpanel is, and where it is *the wrong call.*

Stated once, so the rest of this page can be read on its own merits.

Richpanel is an AI-native helpdesk: the AI agents that resolve the repeat work plus the helpdesk your humans work in, on a single meter. A CX Manager AI reads your past tickets, help center, and public reviews, then drafts the policies, escalation rules, and quality rubric, so setup is a configuration exercise rather than a consulting engagement. Your CX team owns the test cases, and a quality agent reviews closed conversations and feeds gaps back into policy. Time to value runs in three stages: a 30-minute proof of concept built live on your own data during the demo, a two-week pilot, then a four-week deployment. The de-risking is contractual, a 50% resolution guarantee in 30 days with your money back if we miss. We hold a SOC 2 Type II report with an unqualified opinion and zero exceptions across roughly 90 controls for the May to October 2025 period, plus independent HIPAA and GDPR assessments. Around 3,000 brands run on the platform. Pricing is in [the pricing breakdown](https://www.richpanel.com/learn/richpanel-pricing).

### Where it is the wrong call

- **Voice has to be a single native pane.** We integrate with Aircall, Dialpad, and JustCall. We do not host phone. If voice is your highest-volume or highest-CSAT channel and a single native console is non-negotiable, weigh that first.
- **Your resolving actions live in a custom OMS or need ERP write-back.** Our integration depth is strongest on Shopify and the common subscription stacks. If the write-calls that close your tickets sit in a bespoke order system, that integration is the project. Scope it before you sign, with us or anyone, and treat a vendor who waves it away as a risk.
- **You already run a standalone agent layer well.** If a Decagon or Sierra deployment is live and the setup cost is sunk, switching to save setup time is not a good reason. The question worth asking there is what the second bill for the helpdesk underneath actually costs you.
- **Your queue is genuinely high-judgment and high-AOV.** If the human relationship is the product, keep people on the front line. The capacity model on this page assumes a queue that is mostly repeat work, and yours is not.
- **You need a reference at exactly your scale before you can move.** See the honest note above. Ask, and structure the pilot around what you get back.

## What this approach *cannot do.*

- **It does not do the change management.** Seventy agents is seventy people whose day changes, and rollouts that fail on a large floor usually fail on adoption and trust, not on model quality. Budget for the internal communication, the retraining, and a true answer to "what happens to my role", because your team will ask on day one.
- **It does not fix demand created upstream.** If WISMO is 35% of your queue because fulfilment is late, automation lowers the cost of apologising. The eliminate treatment is first in the framework for this reason, and it is the one that needs a partner outside your function.
- **It cannot resolve what it cannot act on.** Resolution ceilings at your integration depth, not at the model's ability. Every category whose write-call sits in a system the AI cannot reach stays with humans until that integration exists.
- **It does not arrive at full rate on day one.** Expect a calibration period after the test cases are authored and real traffic tunes the policies. A vendor promising day-one performance is selling rather than measuring, and you should hold the guarantee to a dated window instead.
- **It does not retire your quality function.** It changes what that function reviews: less manual sampling of human replies, more policy design, escalation review, and adjudicating what the AI flagged. Different job description, and someone has to own it.

## Five moves, in order, *before you talk to a vendor.*

Every one of these is doable in 30 days with the data you already have, and doing them first changes what a vendor can sell you.

### 01. Tag one full week of contacts by category.

Record volume and average handle time per category. This one artifact is the input to everything else on this page, and almost no team has it in usable form when the vendor calls start.

### 02. Compute your capacity identity from your own numbers.

Handling hours per agent per month, contacts per agent, utilisation, and the real cost of a net new head after ramp and backfill. Now you have a ceiling and a marginal cost, both defensible in a budget conversation.

### 03. Assign one treatment to every category above 3% of volume.

Eliminate, self-serve, automate, assist, or human only. For every category marked automate, name the write-call that ends the conversation and the system it lives in. That list is your integration requirement, and it is worth more in a vendor conversation than any feature checklist.

### 04. Write the seven-control guardrail spec and send it out first.

Before the demos. Vendors respond to a control document very differently than to an open discovery call, and those differences are the most useful signal you will get all quarter. A fuller structure sits in our [vendor RFP template](https://www.richpanel.com/learn/ai-customer-service-vendor-rfp-template).

### 05. Fix the scorecard and the reopen window in writing before the pilot.

Resolution rate by category with the reopen window stated, reopen rate, escalation precision, three-way CSAT, and cost per genuinely resolved conversation. Agreeing these at week zero stops the goalposts moving at week four.

### Seven questions that separate vendors quickly

- Name every write-action the AI can take in my systems, and tell me which of my top ten categories each one resolves.
- What is your resolution rate by category, with the reopen window stated? Not containment, not deflection.
- Who authors the test cases, can my team add our own, and can I see the diff when a policy changes?
- Can I set per-action, per-conversation, and daily aggregate monetary caps myself?
- Am I billed once or twice for one AI-resolved conversation? Model my bill at twice my current volume.
- What happens when my allowance runs out mid-month? Does the AI keep working?
- Give me a reference at my headcount, in my vertical, before the technical evaluation.

## The questions large support teams ask *first.*

### How do you scale support without hiring when you already have 50 to 100 agents?

You lower the share of contacts that need a human rather than raising the number of humans. Sort every ticket category above 3% of volume into one of four treatments: eliminate the cause, move it to self-serve, automate it end to end, or assist the human who handles it. Then hold the automated band to a hard standard, that it can take the write-action that closes the conversation, not just describe it. Hiring is the weakest of the four levers at your scale because ramp time, attrition backfill of 30% to 40% a year, and one more team lead per eight to ten agents consume most of every net new head before it delivers.

### How many tickets can one support agent actually handle in a month?

Work it from handling hours rather than from a headline number. A full-time agent is paid about 173 hours a month. Take off shrinkage of 25% to 35% for holiday, sick time, training, coaching, meetings, and breaks, then apply occupancy of 75% to 85% for the share of floor time actually spent on contacts. That lands near 97 handling hours. At an 8-minute average handle time that is roughly 730 contacts a month; at 12 minutes it is about 485. A 70-agent team therefore absorbs somewhere between roughly 34,000 and 51,000 contacts a month. Substitute your own three inputs before using any of this in a plan.

### What share of tickets can AI realistically resolve at enterprise scale?

It is set by two things: what share of your queue is repeat work, and how many of your resolving actions the AI can actually execute. A mature deployment on a repeat-heavy ecommerce or subscription queue lands in the 50% to 80% band, and 50% in 30 days is the number we guarantee contractually with money back if we miss. But that ceiling is capped by integration depth, not by the model. If the calls that close your tickets live in a custom order system the AI cannot reach, those categories stay with humans regardless of how good the AI is. Ask any vendor for resolution rate by category with a reopen window stated, because the average hides which bands are working.

### Should we run AI as a copilot first, or go straight to autonomous?

Copilot first, per category, and treat it as a control rather than a timidity phase. The AI drafts, your agents approve, you watch accuracy on your real traffic by category, and you release autonomy where the numbers hold. You get the throughput gain during that period anyway, and you get an evidence base instead of a leap. The failure this prevents is the one everybody has heard about: a blind cutover, one bad week, and an AI programme that gets switched off permanently for a reason that was a configuration problem.

### How do we stop an AI agent from issuing a refund it should not?

With controls, not with confidence. Identity verification before any account action. Monetary ceilings at three levels: per action, per conversation, and per day in aggregate, because the risk at your volume is not one large refund but a policy misreading applied at scale for hours. An enumerated action inventory, so the AI can only call tools you enabled with arguments you allow, rather than composing operations against production. Explicit escalation when it is out of scope. A full audit log with reversal where the underlying system allows it. And test cases your team owns, run before any policy change ships. Ask a vendor to demonstrate all six, and ask to see three real conversations where the AI stopped and escalated.

### What should we measure in an AI support pilot?

Seven numbers: resolution rate by category with the reopen window stated, reopen rate on AI-closed conversations within 7 days, escalation rate paired with escalation precision, CSAT split three ways (AI-resolved, human-resolved, and human-after-AI), cost per genuinely resolved conversation including the human minutes spent on AI-touched conversations, capacity released in agent hours, and guardrail events caught before send versus reached the customer. Keep containment rate, deflection rate, AI-touched share, and first response time off the success criteria. All four can improve while your queue grows.
