AI Voice Agent Metrics: Containment, FCR, AHT & CSAT Explained

Shambhavi Sinha
View Author Profile
Featured
AI & Solutions
September 23, 2026

Table of contents

Summarize blog with

AI voice agent metrics measure how well an automated voice agent resolves customer conversations, how efficiently it handles them, and how satisfied callers are afterwards. The four that matter most are containment (the share of calls the agent resolves without a human), first-contact resolution (whether the issue was actually fixed on the first try), average handle time (how long a resolved interaction takes), and CSAT (how the customer rated the experience). Read together rather than in isolation, they tell you whether an AI voice agent is genuinely working or just deflecting calls.

The trap most teams fall into is optimising one number in a way that quietly damages another. A containment rate of 80% looks like a win until you notice those “contained” callers are phoning back the next day. This guide defines each metric precisely, shows how they interact, gives the published benchmarks that exist (and flags where they don’t), and explains how to instrument all four so the numbers can be trusted.

Why AI Voice Agent Metrics Differ From Traditional Call Center Metrics

Traditional contact center KPIs were built around human agents: occupancy, schedule adherence, shrinkage, cost per agent hour. Those still matter for the human side of a hybrid operation, but they don’t capture what an AI voice agent is doing.

An AI voice agent introduces questions a human-only metric set never had to answer. How many conversations did the agent fully resolve on its own? When it couldn’t, did it hand off cleanly or dump the caller into a queue? Did automating the call actually solve the customer’s problem, or just delay a human touch? The metrics below are the ones that answer those questions, and the reason containment sits at the top is that it’s the metric with no equivalent in a purely human call centre.

Containment Rate: The Headline Metric (and the Most Misread)

Containment rate is the percentage of conversations an AI voice agent resolves end to end without escalating to a human agent. If 1,000 calls reach the voice agent and 720 are handled without a transfer, containment is 72%.

Formula: containment rate = calls resolved without human transfer ÷ total calls reaching the voice agent.

It’s the headline metric because it maps most directly to cost: every contained call is one a human didn’t have to take. But it’s also the most misread, for one reason: containment is not the same as resolution. A call counts as “contained” if it never reached a human, whether that’s because the agent genuinely solved the problem or because the caller gave up in frustration and hung up. Those are opposite outcomes that look identical in a raw containment figure.

That’s why containment should never be read alone. Pair it with repeat-contact rate (how many contained callers came back within, say, seven days) and with CSAT on contained calls specifically. A high containment rate alongside a low repeat-contact rate is a genuine win. A high containment rate alongside a rising repeat-contact rate means the agent is deflecting, not resolving.

Published containment benchmarks, and why to treat them carefully

Containment benchmarks vary enormously by automation type. Conversational AI vendor Teneo publishes indicative ranges of 5–10% for traditional menu-based IVR, 10–15% for NLU-based conversational IVR, and 10–20% for chatbot-grade AI voice agents, rising far higher for agentic systems on well-suited intents. It notes these depend “heavily on call complexity and the capability of your automation.”

Two cautions apply to every published containment figure. First, vendor benchmarks describe the vendor’s own best-case deployments. Second, and more important, containment depends almost entirely on intent mix. Simple data-lookup intents — order status, EMI due dates, appointment changes — can contain at very high rates. Complex disputes and emotional conversations should escalate by design, and a high containment rate on those would be a red flag, not a triumph. A containment number quoted without its intent mix is close to meaningless. Benchmark against your own pre-automation baseline per intent, not against someone else’s blended average.

First-Contact Resolution (FCR): Did the Problem Actually Get Solved?

First-contact resolution is the percentage of issues fully resolved in a single interaction, with no follow-up needed. It’s the metric that keeps containment honest, because it measures the thing containment only implies: whether the customer’s problem is actually gone.

Formula: FCR = eligible contacts resolved on first contact ÷ total eligible contacts.

FCR has always mattered in customer service, and first-contact resolution is one of the strongest predictors of customer satisfaction and loyalty. SQM Group, which has tracked the metric across the industry for years, puts the call centre industry average at 71%, classes 70–79% as good and 80%+ as world-class — a level it reports only about 5% of call centres reach. It also finds wide sector variation, with retail averaging around 77% and insurance 75%, against telco at 56% and tech support at 64%. SQM’s long-running finding is that every one-point improvement in FCR corresponds to roughly a one-point improvement in CSAT, which is the clearest evidence that these metrics move together rather than independently.

In an AI voice agent context FCR takes on an extra role: it’s the check on whether “contained” calls were truly resolved. If containment is 75% but FCR on those contained calls is only 55%, a meaningful chunk of automation is producing callbacks rather than resolutions.

Measuring FCR for a voice agent is harder than it sounds, because “resolved” isn’t always visible at the moment the call ends. The practical approach is a repeat-contact window: if the same customer doesn’t contact you again about the same issue within a defined period (commonly 24 hours to 7 days), the first contact is counted as resolved. That window has to be set deliberately — too short and you overcount resolutions, too long and unrelated new issues pollute the measure.

What good looks like: the industry average gives you a floor, but the sharper target is internal — FCR on automated calls approaching FCR on the equivalent human-handled calls. If the agent resolves as reliably as a human on the same intent, automation is working.

Average Handle Time (AHT): The Efficiency Metric With a Catch

Average handle time is the mean duration of a resolved interaction — for a voice agent, typically the conversation time plus any automated wrap-up. It’s the clearest efficiency metric, and the one most prone to being gamed.

Human-agent AHT benchmarks are well established and vary by sector: Nextiva’s 2026 benchmark guidance cites roughly 3–5 minutes in retail, 4–6 minutes in banking, 6–8 minutes in healthcare, and 7–10 minutes or longer in insurance, against a long-standing general standard of around 6 minutes. Those figures are useful as the baseline an AI voice agent is measured against, not as a target for the agent itself — automated handling of a narrow intent is a different unit of work from a human handling a mixed queue.

The catch is that AHT is only meaningful alongside resolution. Shaving thirty seconds off AHT is worthless — or worse than worthless — if the shorter calls resolve less often and drive callbacks. A voice agent can always lower AHT by ending calls sooner; whether that’s efficiency or corner-cutting depends entirely on what happens to FCR at the same time. Falling AHT with steady or rising FCR is real efficiency; falling AHT with falling FCR is the agent rushing callers off the line.

For hybrid operations, AHT also splits usefully into two questions: how long the AI agent spends on contained calls, and how much handle time on escalated calls the human saves because the agent already gathered context before transferring. A good handoff should reduce human AHT, because the customer doesn’t re-explain from scratch. If it doesn’t, the handoff is losing context.

CSAT: The Metric That Overrules the Others

Customer satisfaction (CSAT) is a direct measure of how satisfied customers are with an interaction, usually captured through a short post-call survey and expressed as the percentage of respondents who rated the experience positively. Where containment, FCR and AHT measure what the system did, CSAT measures how it felt to the person on the other end — and when it conflicts with the operational metrics, it generally wins.

Formula: CSAT = positive responses ÷ total survey responses, where “positive” is usually the top two points on a five-point scale.

On benchmarks, the bar has moved. Nextiva’s guidance puts the old general standard at around 75% and the current expectation at 85% or higher, with retail at 85–90%, healthcare above 85%, and banking and insurance above 80%. Worth noting alongside this: SQM reports the gap between FCR and CSAT scores has widened from about four points in 2013 to about eight points more recently — resolving the issue is necessary for satisfaction but increasingly not sufficient on its own.

A voice agent can post excellent containment, FCR and AHT and still score poorly on CSAT if callers found it frustrating to talk to, if it misunderstood accents or code-switching, or if the escalation path felt like a fight. CSAT is where those experiential failures surface. It’s also the metric most worth segmenting: CSAT on contained calls versus escalated calls tells you whether the automation itself is the problem or whether the handoff is.

For a fuller treatment of satisfaction and other experience measures, see our guide on measuring customer experience with CX metrics.

What good looks like: CSAT on automated calls should be at least on par with CSAT on human-handled calls for the same intent. If customers rate the AI agent materially lower, the containment savings may be borrowing against customer loyalty — a trade that rarely pays off long term.

Benchmark Summary

Published figures for the human-handled baseline against which an AI voice agent is measured. Containment has no equivalent cross-industry standard, for the reasons above.

Metric General benchmark By sector Source
First-contact resolution 71% industry average; 70–79% good; 80%+ world-class (~5% of centres) Retail 77%, insurance 75%, tech support 64%, telco 56% SQM Group
Average handle time ~6 minutes (long-standing general standard) Retail 3–5 min, banking 4–6 min, healthcare 6–8 min, insurance 7–10 min+ Nextiva
CSAT 85%+ current expectation (previously ~75%) Retail 85–90%, healthcare 85%+, banking and insurance 80%+ Nextiva
Containment No cross-industry standard; depends on intent mix Menu IVR 5–10%, conversational IVR 10–15%, AI voice agents 10–20% and higher on well-suited intents Teneo (vendor-published)

Treat every row as orientation, not a target. The only benchmark that reliably means something is your own pre-automation baseline for the same intents.

How the Four Metrics Interact

The single most important idea in this whole topic is that these metrics are a system, not a scorecard. Each one can be improved in isolation in ways that damage the others.

Metric What it measures Fails when read alone because… Read it alongside
Containment Share of calls resolved without a human Counts frustrated hang-ups as “contained” Repeat-contact rate, CSAT
FCR Whether the issue was actually solved Hard to see at call-end; needs a follow-up window Containment, repeat-contact rate
AHT Efficiency / interaction duration Can be lowered by rushing calls FCR
CSAT How the customer felt A number without a reason; needs segmenting FCR, escalated-vs-contained split

The healthy pattern is containment and FCR rising together, AHT falling while FCR holds, and CSAT on automated calls staying level with human-handled calls. Any metric moving in the right direction while a paired metric moves in the wrong one is a signal to look closer, not to celebrate.

How to Instrument These Metrics

Most measurement disputes are really definition disputes. Before reporting any of these four, settle the following in writing so the numbers mean the same thing month to month.

  • Define the denominator for containment. Decide whether calls that never reach the agent (drops before connect, wrong numbers, silent calls) are excluded. Including them deflates containment; excluding them without documenting it makes the figure incomparable later.
  • Separate hang-ups from resolutions. Tag call endings distinctly: task completed, caller ended early, agent ended, transferred. Without this split, containment cannot be trusted at all.
  • Fix the FCR window and the matching rule. Choose the repeat-contact window (24 hours to 7 days is typical) and decide how you match a repeat contact to the original issue — same customer plus same intent tag is the usual approach. Same customer alone counts unrelated calls as failures.
  • Decide what counts inside AHT. State whether queue time, hold, and automated wrap-up are in or out. Comparing an AI agent’s conversation-only AHT against a human AHT that includes after-call work is not a like-for-like comparison.
  • Control CSAT sampling bias. Post-call survey response rates are low and skew towards strong opinions. Keep the survey trigger, channel and question wording identical for automated and human-handled calls, or the comparison between them is meaningless.
  • Tag every call with an intent. This is the one that makes the other five useful. Every metric here should be reportable per intent, because blended averages hide exactly the intents where automation is failing.

Beyond the Big Four: Supporting Metrics Worth Tracking

The four above are the core, but a few supporting measures round out the picture:

  • Repeat-contact rate — the share of resolved contacts that come back within a set window. The single best guard against mistaking deflection for resolution.
  • Escalation rate and handoff quality — how often the agent transfers, and whether the human receives full context. A clean handoff is as important as a high containment rate.
  • Latency — the response delay in the conversation. High latency wrecks CSAT regardless of how well the agent reasons, because callers talk over the agent or assume the line dropped.
  • Automated QA coverage — with AI-driven quality analysis, 100% of conversations can be scored for script adherence and sentiment, versus the 2 to 5% a manual quality management process typically samples. That fuller coverage is what makes the other metrics trustworthy at scale.

How to Improve AI Voice Agent Metrics Without Trading One for Another

The goal is never to maximise a single metric; it’s to move the system without breaking a link in it.

To improve containment without hurting FCR: expand the agent’s coverage on intents it already resolves well before adding complex ones, and make sure the fallback to a human is easy, so “containment” never comes from callers giving up.

To improve FCR: ground the agent in accurate, current data (most failed resolutions trace back to the agent not having the information it needed), and make sure it can complete the actual task, not just describe it.

To improve AHT safely: reduce dead time and automate wrap-up rather than compressing the conversation itself; a good agent assist layer shortens the human side of hybrid calls by surfacing context automatically.

To improve CSAT: fix latency and accent/language handling first, since those are the most common experiential complaints, and segment CSAT by intent so you’re improving the calls that are actually scoring badly rather than an average.

A platform like Exotel’s AI voice agents reports up to 75% containment alongside conversation quality analysis that scores every interaction, which is the combination that lets teams push containment while keeping FCR and CSAT visible rather than flying blind on them. As with any vendor figure, treat that as a ceiling for well-suited intents and validate it on your own call recordings.

Frequently Asked Questions

What is a good containment rate for an AI voice agent?

There’s no universal number, because containment depends heavily on intent mix. Published vendor ranges run from roughly 5–10% for menu-based IVR up to 10–20% for chatbot-grade AI voice agents, with much higher rates on simple data-lookup intents such as order status or payment due dates. Complex disputes should escalate by design. A containment figure is only meaningful when read alongside repeat-contact rate and CSAT, since a high rate driven by frustrated hang-ups is a problem, not a success.

What is a good FCR rate?

SQM Group puts the call centre industry average at 71%, with 70–79% counting as good and 80% or higher as world-class — a level roughly 5% of call centres reach. Sector averages vary widely, from about 77% in retail down to 56% in telco. For an AI voice agent, the more useful target is internal: FCR on automated calls matching FCR on human-handled calls for the same intent.

What is the difference between containment and first-contact resolution?

Containment measures whether a call was handled without reaching a human; first-contact resolution measures whether the customer’s problem was actually solved. A call can be contained without being resolved — for example, if the caller gave up. Reading the two together is what separates genuine automation from mere deflection.

How is average handle time measured for an AI voice agent?

AHT for a voice agent is typically the conversation duration plus any automated wrap-up time, averaged across resolved interactions. Whether queue time, hold and after-call work are included has to be defined explicitly, or comparisons against human AHT are not like-for-like. It should always be read alongside first-contact resolution, because a lower AHT is only a gain if resolution holds steady.

Can an AI voice agent have high containment but low customer satisfaction?

Yes, and it’s a common warning sign. High containment with low CSAT usually means callers are being kept in automation they find frustrating, whether from latency, poor language handling, or a difficult escalation path. It indicates the containment savings are coming at the expense of customer experience.

Which AI voice agent metric matters most?

None in isolation. Containment maps to cost, FCR to whether problems are solved, AHT to efficiency, and CSAT to customer experience, but each can be improved in ways that damage the others. The most useful signal is the relationship between them, particularly containment and FCR rising together while CSAT holds.

How does automated quality analysis relate to these metrics?

AI-driven quality analysis can score 100% of conversations rather than the small sample a manual process reviews. That fuller coverage is what makes containment, FCR and CSAT trustworthy at scale, since the metrics are only as reliable as the share of conversations actually being evaluated behind them.

Found this interesting? Share it now!

Revolutionize Customer Experience

Discover strategies to enhance customer satisfaction with cutting-edge tools.

Request Demo

Shambhavi Sinha explores the evolving world of technology, with a focus on contact centers, artificial intelligence, and customer experience. She delves into industry trends, breaking down complex concepts to provide valuable insights for businesses and professionals. Through her writing, she aims to keep readers informed about the latest innovations shaping the future of customer communication.

Related Articles

When NOT to Use an AI Voice Agent (An Honest Guide)
Blog

When NOT to Use an AI Voice Agent (An Honest Guide)

AI Call Assistant: What It Is & How Businesses Use It
Blog

AI Call Assistant: What It Is & How Businesses Use It

Voice Streaming API for Telephony: What Enterprises Need Beyond Realtime Models
Blog

Voice Streaming API for Telephony: What Enterprises Need Beyond Realtime Models