AI Contact Center Metrics: KPIs That Survive Scrutiny

AI Contact Center Metrics: KPIs That Survive Scrutiny

ON THIS PAGE

Written by: Matt Beucler, CEO, Plura AI

Key Takeaways

  • Contact center AI performance metrics must separate containment from true resolution. Containment without resolution becomes a vanity metric that fails executive scrutiny.
  • The AI performance hierarchy spans seven layers: interaction, understanding, correct answer, resolution, satisfaction, repeat contact, and cost per resolution.
  • Legacy KPIs like AHT, CSAT, and FCR need segmented, outcome-based versions that distinguish AI-only, AI-assisted, escalated, and human-only interactions.
  • Re-contact rate acts as the critical check on false containment, and cost per resolution rewards completed outcomes.
  • Plura AI unifies voice, SMS, RCS, and AI webchat on a single Stateful Conversation Database so every layer of the performance hierarchy becomes measurable and defensible from one platform.

The AI Performance Hierarchy: From Interaction To Cost Per Resolution

The seven-layer AI performance hierarchy tracks the journey from interaction initiated through intent understood, question answered correctly, issue resolved, customer satisfied, no repeat contact, and lower cost per resolution. Each layer has its own metrics and failure modes. A dashboard that scores well at the top of the hierarchy and poorly at the bottom flatters the deployment instead of describing it.

Most vendor dashboards stop at the interaction and understanding layers. They report containment, handle time, and intent recognition without tying those signals to resolution, satisfaction, or repeat-contact behavior. Zendesk notes that legacy KPIs like deflection, containment, and average handle time were built for a world where humans handled one conversation at a time and do not translate cleanly to measuring AI in customer service.4

Plura’s Stateful Conversation Database runs voice, SMS, RCS, and AI webchat on a single data layer. Leaders can measure every layer of this hierarchy from one platform instead of stitching reports together from multiple point tools.

Plura Unified Inbox interface showing centralized AI Voice, SMS, RCS, and Webchat conversations in one omnichannel workspace.
Plura Unified Inbox centralizes AI Voice, SMS, RCS, and Webchat conversations into one streamlined omnichannel communication workspace.

Run your numbers through Plura’s ROI calculator to check your cost per resolution in real time.

3

AI Containment Rate vs. True Resolution Rate: The Extractable Comparison

Containment rate and true resolution rate often get blended together in reporting. They measure different outcomes and produce different numbers on the same deployment. The table below shows how each metric is defined, where it can be inflated, and which sources support the definition so you can see where the numbers diverge.

Metric Formula Trap Source
Containment rate AI-contained interactions ÷ AI interactions, where the denominator (engaged, total, or in-scope interactions) varies by definition and materially changes the result Counts absence of transfer, not resolution Inkeep
True resolution rate The percentage of customer issues fully and permanently resolved, calculated as issues resolved without repeat contact divided by total issues or interactions handled, where resolution is confirmed by tracking whether the customer contacts the organization again about the same issue within a defined window (commonly 7–14 days) Requires defined re-contact window NiCE
Re-contact rate Case-based Repeat Contact Rate = Index cases followed by a repeat within the window ÷ Total eligible cases (distinct from the customer-based formula, which divides customers who re-contacted by total customers who contacted) Window definition changes the number Umbrex
Automated resolution rate Issues fully resolved by AI ÷ total issues handled by AI Excludes abandoned, partial, escalated Zendesk

A contained interaction is one that ends without transfer to a human agent. A resolved interaction is one where the customer’s issue was addressed completely with no repeat contact within a defined window. A call contained today that returns tomorrow for the same intent has not been resolved, and defensible measurement subtracts 7-day re-contact for the same intent from the numerator.

A contact center taking 100,000 calls a month where the voice AI engages on 60,000 and 30,000 finish without escalation produces three different numbers. The vendor quotes 50% containment (30,000 ÷ 60,000 engaged). Operations measures 30% (30,000 ÷ 100,000 total). Finance reports 24% resolved after subtracting the 6,000 callers who rang back within 7 days about the same issue. Three numbers, one deployment. Re-contact rate is the check on false containment.

Defensible containment targets vary by intent complexity: transactional intents 60-80%, mixed service intents 35-55%, complex intents 10-30%, blended enterprise call mix 25-45%. Figures materially above those bands deserve scrutiny of the denominator and measurement method.

Intent Recognition Accuracy and Knowledge Coverage: The Understanding Layer

Containment and resolution describe what happened at the end of an interaction. The understanding layer shows whether the AI correctly identified the customer’s need before responding.

Intent recognition accuracy = (correctly classified intents ÷ total intents) × 100, expressed as a percentage. This rate shows how often the AI correctly understands what a caller or chatter wants before attempting an answer. First contact resolution rates below roughly 60% suggest responses without resolutions.

Knowledge coverage = AI coverage rate = (Number of contact categories the AI can handle autonomously ÷ Total distinct contact categories) × 100. This is a capability measure of the share of contact types the AI can handle. A volume-weighted alternative = (Volume of contacts in AI-capable categories ÷ Total contact volume) × 100. A deployment can show high intent recognition on a narrow configured intent set while the long tail routes to humans or fails silently.

Zendesk recommends categorizing handoff reasons into policy exception, authentication requirement, missing data, low confidence, negative sentiment, VIP or high-risk account, compliance or privacy concern, and repeated failed resolution attempt so escalation data becomes an improvement roadmap instead of a number to minimize without context.

Plura’s conversation intelligence layer surfaces which intents are failing and why. That signal feeds directly into workflow tuning instead of leaving operators to diagnose failures manually.

Plura Conversation Intelligence dashboard displaying AI-powered call analytics, transfer tracking, and customer conversation insights.
Plura Conversation Intelligence gives businesses AI-powered analytics, call transfer tracking, and customer interaction insights across every conversation.

AI Handle Time Segmentation: Why Blending Lets AI Take Credit

Once you know whether the AI understood the intent, the next question is how long each type of interaction takes.

AI handle time = AI Handling Time = Total Duration of AI Interactions ÷ Number of AI Handled Queries, segmented by interaction type. The four segments that matter are AI-only resolutions, human-only resolutions, AI-assisted human resolutions, and AI escalations requiring agent follow-up.

Without AHT segmentation, AI may reduce the time agents spend on routine requests while increasing the complexity of conversations that reach humans. A higher human AHT can therefore mean the AI is filtering routine work correctly. A blended AHT that falls after AI deployment may simply reflect that easy interactions left the human queue, not that the AI improved efficiency.

ML6 recommends measuring AI-assisted AHT reduction against a baseline of human-only calls. A realistic target is 25% to 35% reduction, with roughly 30% AHT reduction typical in ML6’s voice AI deployments.3 The reduction comes from the AI handling the first two minutes of intake, authentication, and intent capture before the agent picks up.

MessageMind recommends reporting CSAT as three separate lines, AI-only, AI-assisted human, and escalated interactions, rather than one blended number, because the split reveals whether AI is lifting or dragging satisfaction. The same segmentation logic applies to handle time.

Automated Resolution Rate: Measuring Agentic AI Outcomes

Automated resolution rate = issues fully resolved by AI ÷ total issues handled by AI. This metric applies specifically to agentic AI connected to backend systems that complete transactions end-to-end without human involvement.

A true automated resolution requires the AI to deliver an accurate answer, complete any required action, and not require follow-up. Automated resolution rate should exclude abandoned conversations, partial answers, escalations, generic replies, unresolved closures, contained interactions where the user gives up, and repeat contacts for the same issue.

The trap here is counting the end of a conversation as an outcome. A session can end because the answer landed, the customer gave up, the browser closed, or the timeout fired. Only the first of those four outcomes is a resolution. Unlike response-based metrics that measure activity, automated resolution rate measures outcomes: how often automation actually resolves the problem.

Plura’s no-code workflow builder connects AI agents to backend systems so resolution requires an action completed, not just an answer delivered.

Plura Managed Workflows interface showing AI conversation workflows, automation logic, scripts, and operational process management.
Plura Managed Workflows gives businesses fully built AI conversation workflows designed to automate customer engagement and operational tasks.

AI Quality and Hallucination Measurement: Sampling and Groundedness

Hallucination rate = unsupported verifiable claims ÷ all verifiable claims, evaluated under a declared evidence contract that specifies which sources the application was permitted to use.

The sampling methodology: randomly sample 2-5% of live AI outputs and route them to human reviewers for factual verification. At 10,000 queries per day with 2% sampling, this yields 200 reviewed outputs per day, enough to produce a statistically meaningful weekly hallucination rate.

Groundedness checks use NLI (natural-language inference) models to check entailment between each claim and its supporting passage, flagging claims that are neutral or contradicted. A hallucination rate belongs to a full configuration, including model version, prompt, decoding temperature, chunk size, retrieval depth, and reranker, not to a model name.

The denominator trap: a single hallucination percentage is misleading because the denominator is ambiguous. A 3% per-claim rate on an agent making 40 claims per task means almost no task is clean, while a 3% per-task rate means 97 of 100 jobs are trustworthy. Because the two rates tell different stories, report both alongside the average number of claims per response.

Inferred CSAT vs. Survey CSAT: Reconciling Two Signals

Inferred CSAT is an AI-generated customer satisfaction score from 1 (least satisfied) to 5 (most satisfied), predicted entirely from conversation transcripts in real time without any user input.

Survey CSAT is customer-reported satisfaction collected via post-interaction survey. These surveys typically capture responses from roughly 5–20% of interactions, with post-interaction support surveys sometimes reaching 10–25%.

Sentiment modeling can misread neutral or transactional interactions as positive and can miss dissatisfaction that customers do not express in the conversation. Zendesk recommends evaluating customer experience quality beyond surveys using AI-inferred CSAT, sentiment analysis, conversation reviews, QA scoring, and repeat contact data, because surveys only capture the slice of customers who respond.

The reconciliation approach reports inferred CSAT and survey CSAT side by side and investigates divergence. Reporting CSAT as three separate lines, AI-only, AI-assisted human, and escalated interactions, rather than one blended number, reveals whether AI is lifting or dragging satisfaction.

Book a live demo with Plura to see how contact center AI performance metrics surface natively across every channel.

Plura Webchat interface showing AI-powered customer messaging, automated responses, and real-time conversational engagement.
Plura Webchat delivers AI-powered customer conversations with real-time engagement, automated responses, and seamless appointment scheduling.

The Executive Scorecard: Customer, AI, Operations, and Financial Layers

A VP of Contact Center Operations needs a scorecard organized by four layers, not a single blended dashboard number. The customer and AI layers show whether the deployment works for customers. The operations and financial layers show what it costs to make that performance sustainable. Read together, these layers prevent a strong AI layer from masking a weak financial outcome.

Customer layer:

  • CSAT segmented by AI-only, AI-assisted, escalated, and human-only
  • Customer effort score
  • Repeat contact rate

AI layer:

  • Containment rate
  • True resolution rate
  • Automated resolution rate
  • Intent recognition accuracy
  • Knowledge coverage
  • Hallucination rate

Operations layer:

  • AI handle time segmented by interaction type
  • Escalation rate
  • Agent assist adoption
  • Human minutes per resolved contact

Financial layer:

  • Cost per resolution

Cost per resolution is the financial metric that survives scrutiny because it rewards completed outcomes. Zendesk defines it as total support cost divided by successfully resolved inquiries, which is stronger than cost per contact. The reason is that a high-deflection AI system can look cheap on a per-contact basis while quietly generating many repeat contacts.

Plura’s AI Conversation Intelligence surfaces outcome-based metrics such as conversion lift, contact rates, and cost per completed action instead of dashboard summaries. Operators get the data layer this scorecard requires.

Legacy KPIs in the AI Era: A Translation Guide

Human-era KPIs break when an AI handles the interaction. Each one needs a direct translation into an outcome-based metric.

  • AHT becomes segmented AI handle time across AI-only, AI-assisted, escalated, and human-only interactions. Blending them lets AI take credit for removing easy interactions from the human queue.
  • CSAT becomes inferred CSAT plus survey CSAT, segmented by interaction type. Blended CSAT misses non-respondents and hides whether AI is lifting or dragging satisfaction.
  • FCR becomes true resolution rate with a re-contact check. FCR counts premature closures; true resolution rate tracks completed outcomes.
  • SLA becomes time to first AI-powered contact. Traditional SLA measures human queue time and ignores AI response time.
  • 80/20 rule becomes intent coverage and knowledge coverage, the same coverage concepts defined in the understanding layer. The long tail of unconfigured intents is where AI deployments fail silently.
  • Productivity measures becomes cost per resolution instead of cost per contact. Cost per resolution aligns spend with completed outcomes.

Frequently Asked Questions

What Metrics Are Used to Measure AI Performance?

The core contact center AI performance metrics are containment rate, true resolution rate, automated resolution rate, intent recognition accuracy, knowledge coverage, segmented AI handle time, hallucination rate, inferred CSAT, re-contact rate, and cost per resolution. These metrics sit across a hierarchy from interaction through understanding, resolution, satisfaction, repeat contact, and cost. The hierarchy, not any single metric, provides the full picture.

How Do You Calculate Containment Rate and True Resolution Rate?

Containment rate is AI-contained interactions divided by AI interactions, as defined in the comparison table above. True resolution rate is the share of issues resolved without repeat contact within a defined window. The gap between these two numbers on the same deployment represents the difference between what vendors report and what finance approves.

What Is a Good AI Containment Rate?

There is no universal benchmark. The intent-complexity bands in the comparison section above are the defensible ranges, and figures materially above them deserve scrutiny of the denominator. The only defensible headline number subtracts 7-day re-contact for the same intent from the numerator.

How Is AI Handle Time Different from Human AHT?

AI handle time must be segmented across the four interaction types described above. The blended figure is misleading for the reasons already covered: it can fall simply because the mix of interactions reaching humans changed. Segmentation is the only way to read the signal correctly.

What Is Inferred CSAT and Can You Trust It?

Inferred CSAT is the transcript-based score defined above. It should be reported alongside survey CSAT, not as a replacement, because sentiment modeling can misread neutral interactions and miss unexpressed dissatisfaction. When the two signals diverge, that divergence becomes a diagnostic input, especially when CSAT is segmented by AI-only, AI-assisted, and escalated interactions.

Conclusion: Measure What Survives Scrutiny

The AI performance hierarchy, from interaction through understanding, resolution, satisfaction, repeat contact, and cost, provides the organizing logic that separates defensible measurement from flattering dashboards. Containment without true resolution becomes theater. The metrics that survive executive scrutiny are tied to resolved outcomes and repeat-contact behavior. Re-contact rate checks false containment, and cost per resolution rewards completed outcomes.

Plura AI is built for this measurement standard. Plura owns its FCC-licensed carrier, runs voice, SMS, RCS, and AI webchat on one Stateful Conversation Database, and supports TCPA compliance, DNC compliance, HIPAA, SOC 2, ISO certification, GDPR, and SHAKEN/STIR caller ID verification, with compliance features enforced inside the platform on outbound contacts.1,2 Every layer of the AI performance hierarchy is measurable from one platform instead of assembled from disconnected point tools.

See how your cost per resolution compares with Plura’s ROI calculator. Compare contact center AI plans and rates side by side.


1 Plura AI maintains SOC 2, HIPAA, ISO, and GDPR posture as part of its platform infrastructure. References to compliance frameworks in this article describe Plura’s platform capabilities and do not constitute a guarantee that any customer using Plura will themselves be compliant with applicable laws or standards. Customers remain solely responsible for their own regulatory obligations, certifications, consent management, recordkeeping, and the claims they make to their own end users. Consult qualified legal counsel for guidance specific to your use case.

2 This article describes regulatory frameworks at a general level and does not constitute legal advice. Laws and regulations vary by jurisdiction, change over time, and apply differently depending on facts and circumstances. Readers should consult qualified legal counsel before making compliance decisions.

3 Performance figures, customer outcomes, and industry statistics referenced in this article are drawn from cited third-party sources or Plura customer case studies. Individual results vary based on implementation, use case, industry, audience, and execution. Past or aggregate performance is not a guarantee of future results.

4 References to third-party products, services, companies, or research are made for informational and comparative purposes only. Plura AI is not affiliated with, endorsed by, or sponsored by any third party named in this article unless explicitly stated. Trademarks and product names referenced remain the property of their respective owners.

This article is provided for informational purposes only and reflects Plura AI’s understanding at the time of publication. Product capabilities, integrations, and specifications are subject to change. For the most current information, visit plura.ai.

This article was produced with the assistance of AI tools and reviewed by Plura AI prior to publication.

Read Next

See how Plura AI transforms AI voice agents