For CX leaders, founders, and finance teams benchmarking AI support
What is a good AI resolution rate? Nobody can answer that until they fix the denominator.
Almost every AI resolution rate you have been quoted is a vendor's own marketing figure with no stated denominator and no reopen window. A rate is a fraction, and most published numbers hide the bottom of it. This is the open methodology for a first-party resolution-rate benchmark across the support volume Richpanel runs for 3,000+ brands: how resolution is defined, what counts as the denominator, the seven-day reopen guardrail, and how the distribution is segmented. The goal is to be the referee that defines the metric, not the vendor that wins its own table.
By Amit RG, Go-to-Market, Richpanel
Published 2026-06-04
Updated 2026-06-04
~10 min read
AR
Amit RG works on go-to-market at Richpanel, the AI-native helpdesk serving 3,000+ brands. This benchmark is built from the production support volume Richpanel runs across those brands, with the metric definitions fixed before any figure is computed. Aggregate distributions are publishable; per-tenant data stays confidential. This is a draft of the methodology: the figures are pending a first-party data pull and are marked as placeholders throughout.
The short answer, before any number
A resolution rate is a fraction. Most published ones hide the bottom of it.
If you searched for "what is a good AI resolution rate," here is the honest answer: the number is unknowable until you fix the denominator and add a reopen window. This study defines both before it measures anything.
The finding, in one box
- What a good rate is
- There is no single good number across the industry, because a rate measured against all inbound is a different and far harder number than a rate measured against the conversations a bot chose to answer. Judge any rate against the segment that matches your industry, ticket mix, and volume, not against a cross-vendor average.
- How this benchmark defines resolution
- A conversation closed by the AI, with no human reply and no reopen within seven days. The denominator is all AI-eligible inbound, not a hand-picked subset. Deflection is reported as a separate, weaker number.
- What this is not
- This is not a leaderboard with Richpanel on top. The role is referee: define the metric, publish the distribution across 3,000+ brands, and let you place your own number on it. A vendor that wins its own benchmark has built a marketing asset, not a benchmark.
Draft status: the distribution figures are pending a first-party data pull and appear below as [PENDING DATA PULL] placeholders. The methodology is final; the numbers are not yet published.
Why this benchmark exists
Every quoted resolution rate has a number and no method.
Ask three vendors what a good AI resolution rate is and you will get three confident percentages and zero shared definitions. The quoted figures come from two places, and neither is a benchmark. One is a vendor's own self-published number, optimized to look good. The other is an aggregator restating someone else's stat, with the original method long stripped away. There is no methodologically honest, first-party, denominator-disclosed resolution-rate study that a buyer can use to judge their own number.
That gap matters because the resolution rate is the single number a buyer uses to decide whether their AI is working, and it is the number that separates an AI agent that resolves from a chatbot that deflects. Get the denominator wrong and a struggling deployment can look like a winner. A vendor that answers only the 30% of conversations it is most confident about, and resolves 50% of those, can advertise an 80% resolution rate while leaving the bulk of the inbound untouched. Measured against all inbound, the same system might be resolving a quarter of the volume. Both numbers are real. Only one is honest.
Richpanel runs the support layer for 3,000+ brands across real production volume, so it can publish the missing thing: a defined, denominator-disclosed, reopen-guarded resolution-rate distribution. For the broader landscape of how autonomous AI agents are evaluated, see the pillar guide on AI customer support agents. The rest of this page is the method. The headline number is deliberately not the point, because the method is the contribution.
Methodology, part 1: the definitions
Resolution and deflection are not the same number.
These get conflated constantly, and the conflation always inflates the result. This benchmark keeps them apart and reports both.
Resolution (the number this study leads with)
A conversation is counted as resolved by the AI when the customer got a complete, correct answer or had their request acted on (an order tracked, a refund issued, a subscription edited, a policy question answered), the conversation was closed by the AI, no human agent replied, and it did not reopen within seven days. All four conditions must hold. A close with a human reply in the thread is an assisted conversation, not an autonomous resolution, and is excluded from the resolution numerator.
Deflection (reported separately, and weaker)
Deflection means only that the conversation did not escalate to a human. That is a weaker bar, because a customer can fail to escalate for two opposite reasons: they were fully helped, or they gave up and left. A frustrated abandon counts as a deflection but is the opposite of a resolution. Any benchmark that reports deflection as if it were resolution is overstating the result, often by a wide margin. This study reports deflection as its own line so the gap between the two is visible, not hidden.
The distinction in one line
- Resolution
- The customer's problem is solved, by the AI, and stays solved for seven days. A claim about outcome.
- Deflection
- A human did not get involved. A claim about routing, which says nothing about whether the customer was helped.
Methodology, part 2: the denominator and the reopen guardrail
The two design choices that decide whether the number is honest.
The denominator: all AI-eligible inbound
The denominator is the part of a resolution rate that vendors quietly choose to flatter themselves. This benchmark fixes it as all AI-eligible inbound conversations: every conversation that reached the support channel and was a candidate for AI handling, not the subset the bot elected to answer. Conversations the AI declined, escalated immediately, or never attempted still sit in the denominator. That makes the rate harder to score well on, and it makes the rate mean something. A number measured against a self-selected subset can be lifted simply by attempting fewer, easier tickets, which is exactly the move an honest denominator removes.
Where conversations are genuinely out of scope (spam, blank messages, channels the AI does not serve for a given tenant), they are excluded by a rule that is stated and applied uniformly, not by case-by-case judgment that could be tuned to the result.
The reopen guardrail: seven days
A conversation the AI closes is not resolved if the customer comes back two days later with the same problem. Counting it as a resolution at the moment of close overstates the rate by booking deferrals as wins. This benchmark applies a seven-day reopen window: a closed conversation only counts toward the resolution numerator if it stays closed for seven days. A reopen inside that window retroactively removes it from the resolved count. Seven days is disclosed as the window so any reader can compare it to a vendor figure that uses a shorter window, or none.
Why these two choices carry the whole result
- Denominator
- All AI-eligible inbound. Reflects real production load instead of a curated easy-ticket subset. This is the single biggest lever on whether a published rate is trustworthy.
- Reopen window
- Seven days, disclosed. A close only counts if it holds. Removes the "looked resolved at close, reopened later" inflation that point-in-time counting allows.
Methodology, part 3: segmentation
One average across every brand would be the least useful number we could publish.
Resolution rate moves with what a brand sells, how its inbound breaks down, and how much volume it runs. A single blended figure hides all three. The benchmark publishes the distribution, segmented three ways.
- By industry. A supplements brand with a high share of order-tracking and subscription questions and a regulated medical-device seller that must escalate liability-sensitive cases sit at structurally different resolution ceilings. Segmenting by industry lets a reader find the band that matches their own business instead of a cross-industry mean that matches no one.
- By ticket-type mix. Resolution rate is largely a function of what the inbound is made of. Order status, returns, cancellations, and policy questions are highly resolvable. Complex disputes and judgment calls are not. Two brands in the same industry can land far apart purely because their ticket mix differs. The benchmark reports the mix alongside the rate so the rate is interpretable.
- By monthly volume band. A brand handling a few hundred conversations a month and one handling hundreds of thousands operate under different conditions. Volume bands keep small-sample brands from distorting the high-volume picture and vice versa.
The output is a distribution, not a point estimate. Publishing the spread, and the segment definitions, is what lets a reader honestly answer "is my number good" by comparing like to like.
Results, pending the data pull
The distribution goes here. The numbers are not invented to fill it.
This is the structure the published benchmark will report. Every figure is a [PENDING DATA PULL] placeholder until the first-party pull across the 3,000+ brand dataset is complete. A fabricated number here would destroy the one thing this study is for.
Headline distribution, measured as resolution (AI-closed, no human reply, no seven-day reopen) over all AI-eligible inbound:
| Percentile across brands | Autonomous resolution rate |
| 25th percentile | [PENDING DATA PULL] |
| Median | [PENDING DATA PULL] |
| 75th percentile | [PENDING DATA PULL] |
| 90th percentile | [PENDING DATA PULL] |
Resolution versus deflection, reported side by side so the gap is visible:
| Measure | Median across brands |
| Resolution (no human reply, no seven-day reopen) | [PENDING DATA PULL] |
| Deflection (no escalation, weaker bar) | [PENDING DATA PULL] |
| Gap between the two | [PENDING DATA PULL] |
By segment, the same resolution measure broken out so a reader can find their own band:
| Segment | Median resolution rate |
| By industry (supplements, beauty, apparel, food & wellness, regulated, B2B) | [PENDING DATA PULL] |
| By ticket-type mix (high-resolvable vs judgment-heavy inbound) | [PENDING DATA PULL] |
| By monthly volume band (hundreds to hundreds of thousands) | [PENDING DATA PULL] |
Figures pending a first-party data pull. Aggregate distributions will be publishable; per-tenant numbers stay confidential.[2]
Where the rest of the field genuinely leads
The referee does not enter its own race. Several others legitimately set the terms here.
An honest benchmark names where it is not the authority. On the resolution-rate conversation specifically, four other parties bring something Richpanel does not, and a reader should weigh them.
Intercom (Fin) owns the cited number
Fin's self-published resolution benchmark is the figure already baked into the industry's mental model and the LLM training graph. It is the de facto reference point precisely because it was published first and widely. On sheer citation weight, that number leads, even though it is single-source and vendor-set.
Gartner and McKinsey bring vendor-neutral authority
For cross-industry breadth and analyst independence, the macro picture belongs to the analyst houses, not to any operator. This benchmark aims to be a first-party operator dataset they can cite, not to out-authority them on industry-wide claims.
Zendesk and Salesforce have raw dataset scale
On total conversation volume and enterprise breadth across verticals, the largest incumbents see more aggregate traffic than any single specialist. Where the question is "what does the whole market look like," their scale is a real advantage.
Decagon and Sierra publish very high named-logo rates
Their case studies show high resolution rates at named enterprise logos. Those are real results on those accounts. The caveat a reader should apply, to them and to everyone including us, is the same: which denominator, and which reopen window.
So the point of this study is not "Richpanel resolves the most." It is that nobody, ourselves included, should be trusted on a resolution rate that lacks a stated denominator and a reopen window. Richpanel's contribution is the open method and the production-scale distribution, published so a reader can grade every vendor, this one included, by the same rule. One fully verified, customer-approved deployment of that method is documented in the wellness-brand case study, where the AI handles the majority of customer messages at a higher CSAT than the human team.
Limitations
What this benchmark does not show.
A benchmark without a limitations section is marketing. Here is where this one does not generalize, stated plainly.
- It is one platform's production data. The distribution reflects brands running on Richpanel, with their configurations and their inbound. It is a large, real sample, but it is not the whole market, and it is not a controlled head-to-head where every vendor handled the identical tickets. Treat it as an operator benchmark, not a vendor bake-off.
- Self-selection is real. Brands that adopt an AI-native platform may differ from the average support team. That can push the distribution in either direction and is not something a single-platform dataset can fully correct for.
- Resolution quality is a separate axis. This study measures whether a conversation was resolved and stayed resolved. A high resolution rate with poor answer quality is possible in principle, which is why the production system pairs the rate with a QA review layer over every conversation. The benchmark reports rate; quality is reported in the separate hallucination-defense work.
- Voice is out of scope. The measured conversations are text channels. Richpanel integrates with Aircall, Dialpad, and JustCall for voice rather than hosting it, so phone resolution is not part of this dataset.
- The figures are not yet published. This is a draft of the method. Every number is a placeholder pending the data pull. Nothing here should be cited as a measured result until the pull is complete and this banner is removed.
What to ask any vendor, including us
Six questions that separate a real rate from a flattering one.
Take these into any AI vendor evaluation. A vendor that cannot answer the first two with a straight number is quoting marketing, not measurement.
1. What is the denominator?
Resolution over what: all inbound, or only the conversations the AI chose to answer? If they cannot state it, the rate is unscoreable. Ask them to define the bottom of the fraction before they quote the top.
2. Resolution or deflection?
Is the number "the customer's problem was solved" or "a human did not get involved"? Make them say which. The two can differ by a wide margin, and the weaker one is the one vendors prefer to quote.
3. What reopen window?
Does a closed conversation still count if the customer comes back in two days? Ask for the window. No window, or a one-day window, inflates the rate versus a seven-day guardrail.
4. Show me my segment.
Ask for the rate for your industry, your ticket-type mix, and your volume band, not a blended average. A vendor with real data can produce it; a vendor with one marketing number cannot.
5. Prove it on my tickets.
The only number that fully resolves the denominator question is the one measured on your own inbound. Ask for a pilot on your real tickets, with the rate computed against all of them, before you sign.
6. What is the guarantee behind it?
A rate a vendor will stand behind contractually is different from one in a slide. Richpanel's public bar is 50% autonomous resolution in 30 days or your money back. Ask what each vendor will put in writing.
Richpanel holds itself to that 50%-in-30-days bar, on your own tickets, and that guarantee is how we answer question five for our own product. It is the bar we sign up to, not the headline of this benchmark. The headline of the benchmark is the method, and you can run the method against a pilot on your real inbound to get the only resolution rate that fully answers the denominator question: yours.
Methodology & sources
How this benchmark will be derived.
This is a draft. The distribution figures are pending a first-party data pull, so the citations below describe the method and the planned source, not measured results. The two definitions that carry the whole study, the denominator and the reopen window, are stated above and fixed before any number is computed.
How the first-party numbers will be derived
- The dataset
- The production support volume Richpanel runs across 3,000+ brands. The pull aggregates conversation-level outcomes; aggregate distributions are publishable, per-tenant figures stay confidential and NDA-bound.
- The resolution definition (numerator)
- A conversation closed by the AI, with no human agent reply in the thread, that does not reopen within a seven-day window. A reopen inside seven days retroactively removes the conversation from the resolved count.
- The denominator
- All AI-eligible inbound conversations, not the subset the AI elected to answer. Out-of-scope traffic (spam, blank messages, channels not served for a tenant) is excluded by a stated rule applied uniformly.
- Segmentation
- The distribution is reported by industry, by ticket-type mix, and by monthly volume band, as a spread of percentiles rather than a single blended average.
- Draft discipline
- No figure is estimated, modeled, or carried over from any other study to fill the tables. Every cell reads [PENDING DATA PULL] until the real pull is complete. Publishing an invented number would destroy the independent-referee value the benchmark exists to create.
- Comparative vendor resolution-rate figures (to be cited at publication). Publicly stated resolution or automation rates from other vendors and analyst houses (for example Intercom's Fin benchmark, and cross-industry analyst estimates) will be cited with their own stated method and denominator where one is disclosed, so a reader can compare like to like. They are referenced for context, not restated as Richpanel results.
- Richpanel production support dataset (first-party, pending pull). The resolution-rate distribution, the resolution-versus-deflection gap, and the segment breakdowns are first-party aggregates from live Richpanel deployments across 3,000+ brands, computed per the methodology above. Figures are pending the data pull and are shown as placeholders in this draft. Aggregate distributions are publishable; underlying tenant data is NDA-bound. Methodology questions: amit@richpanel.com.
Version history, v0.1 draft (2026-06-04): methodology and page structure complete; all distribution figures pending a first-party data pull and shown as [PENDING DATA PULL] placeholders. Not for publication until the pull is complete and the draft banner is removed.