Hero Background

Power Your Software Testing with AI Agents and Cloud

The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.

AIAI TestingAgent Testing

Chatbot Metrics: 15 KPIs to Measure Chatbot Performance

Chatbot metrics explained: 15 KPIs with formulas, the unit each one is counted in, what skews containment, resolution and CSAT, and what to test before release.

Published on:

OVERVIEW

On November 11, 2022, the day their grandmother died, Jake Moffatt asked the chatbot on Air Canada's website about bereavement fares. The tribunal decision in Moffatt v. Air Canada records the chatbot's reply: anyone who had already traveled could submit the ticket for a reduced bereavement rate within 90 days of the date it was issued. The airline's own bereavement travel page said the policy did not apply to requests made after travel was completed.

By two of the usual chatbot metrics, a reply like that is a success: it is not a fallback, and a chat that ends without a transfer to a person counts as contained. On February 14, 2024, British Columbia's Civil Resolution Tribunal found that the airline "did not take reasonable care to ensure its chatbot was accurate" and ordered it to pay $812.02 in Canadian dollars for damages, interest and fees.

Every number on a chatbot analytics dashboard depends on which events count and what they are divided by. Each of the chatbot KPIs in this guide comes with its formula, the unit it is counted in, what skews it and what you can test before release. The chatbot testing guide covers the testing process itself.

Overview

Chatbot metrics are measurements, mostly ratios, computed from conversation logs, labels and survey responses, that show whether a chatbot understood the user, resolved the request and answered correctly. Track outcome, handoff, understanding, experience and answer quality metrics together, and state the formula and the denominator behind every number you report.

How do you measure chatbot performance?

  • Goal completion rate: Goal completion rate is the share of engaged chatbot sessions in which the user's goal was achieved, confirmed by a back-end event or a labeled outcome. Resolution rate and containment rate should both be checked against goal completion rate.
  • Escalation rate: Escalation rate is the share of engaged chatbot sessions handed to a person, and containment rate is its complement. Report escalation rate by reason, because a handoff required by a business rule is the chatbot working as designed and a handoff after repeated failures is not.
  • Fallback rate: Fallback rate is the share of bot replies that returned a fallback or no-match message. The same chatbot log gives one fallback rate per reply and another per conversation, so a report has to name the unit it used.
  • CSAT: CSAT (customer satisfaction score) is the average rating chatbot users give in a post-chat survey. The score covers only the users who answered the survey, so report the response rate with every CSAT figure.
  • Hallucination rate: Hallucination rate is the share of checked chatbot responses containing at least one claim that contradicts the source or cannot be verified from it. Groundedness scores the same check claim by claim, as supported claims divided by all claims in a response.

How do you measure hallucinations in an LLM chatbot?

Fix a question set with a reference for each answer and run every question several times. Count the share of responses with a contradicted or unverifiable claim (hallucination rate) and the share of claims the source supports (groundedness). TestMu AI's Agent Testing scores hallucination detection on simulated chat conversations before a release.

What Are Chatbot Metrics?

Chatbot metrics are measurements, mostly ratios, that describe how a chatbot's conversations went: whether the user's goal was met, whether a person had to take over, whether the bot understood each message, what the experience was like and what it cost, and, for LLM chatbots, whether the answers were true to their sources and to the conversation so far.

Every metric is built from the same parts:

  • Event definition - what counts as a fallback, a handoff, a resolution or a completed goal.
  • Numerator - how many of those events happened in the period.
  • Denominator and unit - what the events are divided by, counted in a stated unit such as a session, a bot reply or a claim.

Two dashboards can print different values from the same conversations because they chose different events or units.

The groups in this guide follow a model from 1997. PARADISE, a framework for evaluating spoken dialogue agents by Walker, Litman, Kamm and Abella, models performance as "a weighted function of a task-based success measure and dialogue-based cost measures, where weights are computed by correlating user satisfaction with performance."

Chatbot performance metrics still divide that way. Outcome metrics measure task success; handoff, understanding and effort metrics measure dialogue costs; and CSAT measures satisfaction. Chatbots that generate their answers need one more group, answer quality, which asks whether a reply is supported by its source and consistent with earlier turns.

GroupQuestion it answersMetrics in this guide
OutcomeDid the user get what they came for?Goal completion rate, resolution rate, containment rate
Handoff and drop-offWhich conversations left the bot, and why?Escalation rate, abandonment rate, recontact rate
Understanding and effortDid the bot understand, and how much work did the user do?Intent recognition accuracy, fallback rate, average turns to resolution
Experience and costHow did it feel, how fast was it, and what did it cost?CSAT, response latency, cost per resolved conversation
Answer quality (LLM chatbots)Was what the bot said true to its source and to the conversation so far?Hallucination rate, groundedness, context retention

Chatbot Metrics at a Glance

In the table, numbers 1 to 12 describe how a conversation went; 7 and 8 assume an intent model or a defined fallback event. Numbers 13 to 15 check what the bot said, against a source or against earlier turns.

MetricFormulaCounted per
1. Goal completion rateSessions with the user's goal achieved / engaged sessionsSession
2. Resolution rateSessions marked resolved / engaged sessions, reported as confirmed and assumedSession
3. Containment rate(Engaged sessions - escalated sessions) / engaged sessionsSession
4. Escalation rateEscalated sessions / engaged sessions, split by reasonSession
5. Abandonment rateEngaged sessions that ended with no resolution and no handoff / engaged sessionsSession
6. Recontact rateResolved sessions with a repeat contact on the same issue within N days / resolved sessionsResolved session
7. Intent recognition accuracyUtterances given the correct intent / labeled utterancesUtterance
8. Fallback rateFallback or no-match replies / all bot repliesBot reply
9. Average turns to resolutionUser turns in confirmed resolutions / confirmed resolutionsConfirmed resolution
10. CSATSum of ratings / ratings received, shown with the response rateSurvey response
11. Response latencyTime from user message to bot reply, at P50 and P95Bot reply
12. Cost per resolved conversation(Platform fees + model usage + maintenance time) / resolved conversationsResolved conversation
13. Hallucination rateResponses with one or more contradicted or unverifiable claims / responses checkedResponse
14. GroundednessClaims the source supports / all claims in the responseClaim
15. Context retentionTurns that used earlier information correctly / turns that depended on earlier informationTurn

The denominators in the table rest on these terms:

  • Engaged session - a session in which the user asked for something.
  • Unengaged session - a session that opens the widget, says hello and leaves. It stays out of every denominator.
  • Session and conversation - this guide uses both words for the same unit: one continuous exchange between a user and the bot.
  • Volume counts - total sessions, active users and messages per session are not on the list, because they size these denominators and do not say how a conversation went.

Google's Dialogflow CX analytics documentation defines the escalation rate in its intent escalations view as the "Percentage of sessions that requested human escalation" and the no-match rate for each page as the "Percentage of total conversation turns that resulted in a no-match for the page". One analytics panel therefore reports a per-session number and a per-turn number, so write the unit next to every chatbot KPI you publish.

Outcome Metrics: Goal Completion, Resolution, and Containment

Outcome metrics are the chatbot success metrics: they say whether a conversation achieved something. The same word can hide different formulas here, because each platform decides for itself when a conversation counts as resolved or successful.

PlatformIts termWhat countsSource
Microsoft Copilot StudioResolved impliedThe session "is completed without user confirmation but instead based on the agent's logic", for example when the user lets it time out after the End of Conversation topic.Monitor conversational agents
Intercom FinAssumed resolutionAfter Fin's last answer, the customer "exits the conversation without requesting further assistance".Fin AI Agent outcomes
Amazon Lex V2Success (conversation)"The final intent in the conversation is categorized as a success."Key definitions, Lex V2 Developer Guide

Microsoft's outcome model is also an accounting identity. A session is unengaged or engaged, and every engaged session ends as resolved (confirmed or implied, which this guide calls assumed), escalated or abandoned, which gives:

engaged sessions = resolved confirmed + resolved assumed + escalated + abandoned

Containment is everything that was not escalated, so it includes every assumed resolution and every abandoned session. Resolution rate and containment rate are both read off that equation, and goal completion rate is the outside check on them.

1. Goal Completion Rate

Goal completion rate is the share of engaged sessions in which the user's goal was achieved: sessions with the goal achieved / engaged sessions. On a sales or booking chatbot the goal is a meeting booked or an order placed, and the formula stays the same.

  • Data it needs - a goal event for each intent that your own systems can confirm, such as a refund record created or an order status returned, read from the back end where possible and joined to the session ID.
  • What skews it - counting "reached the last step of the flow" as success, which is what a flow-completion rate measures, and defining goals only for the easy intents.
  • Before release - give every test scenario an expected end state and check that state after the conversation, in the system of record where one exists.

2. Resolution Rate

Resolution rate is the share of engaged sessions the platform marks as resolved: sessions marked resolved / engaged sessions. Report it as two numbers: confirmed, where the user said the issue was solved, and assumed, where the bot answered and the user left without saying so.

  • Data it needs - an outcome label on every engaged session, with the confirmed or assumed flag kept instead of merged.
  • What skews it - assumed resolutions, because a user who gave up and a user who was helped look the same in the log; the session timeout, which decides when silence becomes an outcome; and rules such as the Amazon Lex V2 one in the table above, where the conversation takes the result of its final intent.
  • Before release - scenario tests have no assumed resolutions. Each scenario either reached its expected end state or did not, so pre-release resolution is goal completion on the test set.

3. Containment Rate

Chatbot containment rate is the share of engaged sessions that were not handed to a person: (engaged sessions - escalated sessions) / engaged sessions.

  • Other names - if your platform reports a deflection or self-service rate, check its formula. When it is "not escalated / total", it is this metric under another name.
  • Data it needs - a handoff event on the session, including handoffs to phone, email or a ticket form.
  • What skews it - a handoff option that is hard to reach, abandoned sessions and assumed resolutions (both contained by definition), and sessions that never engaged if they are left in the count.
  • Before release - containment is a production number. Test both sides of the handoff decision instead: scenarios that must escalate, such as a request for a person, and scenarios that must not.

Handoff and Drop-Off Metrics: Escalation, Abandonment, and Recontact

Handoff and drop-off metrics describe the conversations that left the bot: to a person, to a timeout, or back again a few days later. Each needs a reason or a time window attached before the number means anything.

4. Escalation Rate

Escalation rate is the share of engaged sessions handed to a person: escalated sessions / engaged sessions. It is the complement of containment. Microsoft Copilot Studio labels each escalated session with its cause:

  • System intended - a business rule set by the bot's maker escalates the session on purpose. Microsoft calls this "an expected outcome of the conversation".
  • System unintended - the session exceeded a threshold set by the maker. "Usually, this outcome indicates that the user is stuck in the conversation and needs assistance."
  • User requested - the user explicitly asked for a person.

A blended escalation rate adds those causes together, so a handoff that worked as designed and a bot that failed look the same in it.

  • Data it needs - the handoff event plus a reason code captured at the moment of escalation.
  • What skews it - the event that counts as an escalation. In Copilot Studio a session is Escalated once the Escalate topic or a Transfer to agent node runs, "whether the conversation transfers to a live agent or not", so the rate can include handoffs that never reached a person.
  • Before release - write a case for each path: a user who asks for a person, a user who fails the same step repeatedly, and a request your policy says must go to a person. Check that the conversation context reaches the person taking over in each one.

5. Abandonment Rate

Abandonment rate is the share of engaged sessions that ended with no resolution and no handoff: abandoned sessions / engaged sessions. The event behind it is a timeout, so the timeout length is part of the definition.

  • Platform examples - Copilot Studio marks an engaged session Abandoned when it "times out after 30 minutes and didn't reach a resolved or escalated state". Dialogflow CX counts abandoned conversations under End Interaction and keeps a separate fail-safe category, Other, for cases such as a user who says hi and immediately closes the conversation.
  • Data it needs - the session end event and the timeout rule that produced it.
  • What skews it - the timeout length, and users who got their answer and closed the window without confirming.
  • Before release - the rate itself needs real users. As a proxy, run impatient and confused user personas against the bot and record the turn at which each one gives up.

6. Recontact Rate

Recontact rate is the share of resolved sessions followed by another contact about the same issue within a set window: resolved sessions with a repeat contact within N days / resolved sessions. It audits assumed resolutions: if the user comes back about the same issue, the earlier session was not resolved, whatever its label said.

  • Platform example - Intercom builds a version of this correction into its own count: when a customer later returns to the same conversation seeking further assistance, the resolution is deducted.
  • Data it needs - a user identity that persists across sessions and channels, an issue or intent key, and a stated window such as 7 days.
  • What skews it - anonymous users, and customers who come back by phone or email, where the chatbot's analytics cannot see them.
  • Before release - there is no direct test. The nearest pre-release signal is completeness: whether the answer gave the user enough to act on.

Understanding and Effort Metrics: Intent Accuracy, Fallback Rate, and Turns

Understanding metrics ask whether the bot recognized what each message meant, and effort metrics ask how much work the user did to be understood. What the log can tell you depends on the kind of chatbot:

  • Intent-based chatbot - matches each message to a predefined intent and emits a no-match event when nothing fits.
  • LLM chatbot - generates every reply and has no such event unless you define one, so its fallback rate is only as meaningful as that definition.

7. Intent Recognition Accuracy

Intent recognition accuracy is the share of user utterances matched to the correct intent: utterances given the correct intent / labeled utterances. A dashboard without labels can only report a match rate.

  • Platform example - Amazon Lex V2 counts an utterance as Detected when it "recognizes the utterance as an attempt to invoke an intent configured for a bot", which says the utterance matched an intent and cannot say whether it was the right one.
  • Data it needs - a labeled utterance set, which means a person decided what each utterance meant. Draw it from real logs and refresh it as products and phrasing change.
  • What skews it - one blended number hiding a failing intent, and clean, invented test phrasings.
  • Before release - compute precision, recall and F1 for each intent on the labeled set. The chatbot testing guide covers building that set.

8. Fallback Rate

Fallback rate is the share of bot replies that were a fallback or no-match message: fallback replies / all bot replies. Counting conversations with at least one fallback gives a different number: 30.6% against 14.2% in the sample log later in this guide.

  • Platform example - Dialogflow CX counts its no-match rate per conversation turn for each page. Its agent settings documentation says a no-match event is invoked when the confidence score for an intent match is below a classification threshold you can tune.
  • Data it needs - a fallback or no-match flag on every bot reply.
  • What skews it - the confidence threshold. Lowering it cuts the fallback rate and raises wrong-intent matches.
  • Before release - send out-of-scope and gibberish inputs and expect a fallback, then count wrong-intent matches as well as fallbacks. The guide to chatbot test cases has fallback and error-handling cases to start from.

9. Average Turns to Resolution

Average turns to resolution is the number of user turns it took to reach a resolved outcome: user turns in confirmed resolutions / confirmed resolutions. PARADISE lists "the number of turns or elapsed time to complete the task" among the efficiency measures it treats as dialogue costs, and for a chatbot the turn count is a direct measure of user effort.

  • Data it needs - turn counts joined to the session outcome.
  • What skews it - averaging over every session, since short abandoned sessions pull the mean down and make the bot look efficient, and including assumed resolutions, which are short because the user left.
  • Before release - set a turn budget for each scenario and fail the scenario when the bot needs more.

Experience and Cost Metrics: CSAT, Latency, and Cost per Resolution

Each metric in this group has a denominator that is easy to get wrong: who answered the survey, which replies were timed, and which conversations count as resolved.

10. CSAT

CSAT (customer satisfaction score) is the average rating users give in a post-chat survey: sum of ratings / ratings received, or the share of ratings that were 4 or 5.

  • Platform example - Copilot Studio reports a score out of 5 averaged over "sessions in which users responded" to the survey, and maps scores of 1 and 2 to Dissatisfied, 3 to Neutral, and 4 and 5 to Satisfied.
  • Data it needs - survey responses joined to sessions, and the number of sessions that were offered the survey.
  • What skews it - who answers. The score covers responders only, and a survey shown at the end of the resolved path never reaches the users who abandoned.
  • Before release - it cannot be tested, because no user has rated anything yet. A satisfaction estimate that a model produces from a transcript is a different measurement, so label it as an estimate.

11. Response Latency

Response latency is the time from the user's message to the bot's reply. Report the 50th and 95th percentiles (P50 and P95) instead of the mean, because a few slow replies disappear into an average.

  • Streaming chatbots - keep two numbers: time to first token, when text starts to appear, and time to the full reply.
  • Data it needs - a timestamp on each user message and each bot reply, taken where the user sees them, plus the first-token time when the bot streams.
  • What skews it - server-side timings that leave out the network and the chat widget, and averages that hide the tail.
  • Before release - run concurrent conversations at your expected peak and read P95. On voice and phone channels, measure speech-to-text accuracy alongside latency; the guide to conversational AI testing covers both.

12. Cost per Resolved Conversation

Cost per resolved conversation is what the chatbot cost to run for a period divided by the conversations it resolved: (platform fees + model usage + maintenance time) / resolved conversations. The numerator covers every conversation, including the ones that escalated or were abandoned, because you paid for those too.

  • Data it needs - the platform invoice, model usage in tokens per reply (a token counter can estimate it from a sample prompt and answer), the hours spent maintaining the bot, and the resolved count.
  • What skews it - which "resolved" sits in the denominator (confirmed, assumed or goal-completed), and dropping the cost of unresolved conversations from the numerator.
  • Before release - measure tokens per test conversation and multiply by expected volume. The resolved count has to come from production.
  • ROI estimate - compare this figure with your own cost per human-handled contact, and use confirmed resolutions: an assumed resolution that turns into a phone call has not replaced one.

Answer Quality Metrics for LLM Chatbots

The first 12 metrics describe how a conversation went, and none of them checks whether what the bot said was true, which was the failure in Moffatt v. Air Canada. Answer quality metrics check the content of a reply against a source (a knowledge base, a policy page or a system of record) or, for context retention, against what was said earlier in the conversation. They matter most for chatbots that generate answers with an LLM, where nobody wrote the reply in advance.

Answer quality scores come from labels instead of logs: a person or a judge model reads each response and decides. A judge model brings its own error. In the 2023 paper Large Language Models are not Fair Evaluators, Wang and colleagues changed only the order in which two answers were shown and found that Vicuna-13B could beat ChatGPT on 66 of 80 test queries with ChatGPT as the evaluator.

Validate any judge against human labels on a sample before you report its rate. The guide to LLM-as-a-judge covers that calibration step.

13. Hallucination Rate

Hallucination rate is the share of checked responses containing at least one claim that contradicts the source or cannot be verified from it: responses with one or more such claims / responses checked. The Survey of Hallucination in Natural Language Generation by Ji and colleagues separates two kinds:

  • Intrinsic hallucination - "The generated output that contradicts the source content".
  • Extrinsic hallucination - output that "cannot be verified from the source content". The survey adds that an extrinsic hallucination "is not always erroneous".

A rate that counts contradictions only will be lower than one that also counts unverifiable claims, so state which kind yours counts.

  • Data it needs - a reference for each question (the policy text, the knowledge-base article or the record) and a label for each response from a person or a judge model.
  • What skews it - one run per question. A generative chatbot can answer the same question differently each time, so run each one several times; the guide to testing non-deterministic AI outputs covers how many runs a rate needs.
  • Before release - build a fact-checked question set and run it before every prompt, model or knowledge-base change. The guide to LLM hallucination detection compares the detection methods.

14. Groundedness

Groundedness is the share of the claims in a response that the source supports: supported claims / all claims in the response. It is scored per claim, so an answer with nine supported claims and one invented claim scores 90% on groundedness and still counts as one hallucinated response under metric 13.

FActScore, published at EMNLP 2023, is a method of this kind: it "breaks a generation into a series of atomic facts and computes the percentage of atomic facts supported by a reliable knowledge source". Its authors report that ChatGPT reached only 58% on biographies of people. That is a 2023 result for one task and one model, and it should not be read as a benchmark for a support chatbot.

  • Data it needs - the source the bot was given for that answer, and the response split into individual claims.
  • What skews it - claim granularity, since one compound sentence can be counted as one claim or three, and scoring only the questions the bot chose to answer.
  • Before release - check each claim against the source. For facts that live in your own systems, such as a price or an order status, compare against the system of record instead of the knowledge base.

15. Context Retention

Context retention is the share of turns that correctly used information given earlier in the conversation: turns that used earlier information correctly / turns that depended on earlier information.

  • Data it needs - multi-turn transcripts with the dependent turns marked. Marking them is the hard part, because someone has to decide which turns depended on something said before.
  • What skews it - short test conversations and single-turn test sets, which contain no dependent turns and so score nothing.
  • Before release - write scenarios that plant a fact early, such as an account number or a delivery address, and need it several turns later, after a change of topic.

How to Measure Chatbot Performance From a Conversation Log

  • Define a goal event for each intent - an event your own systems can confirm, such as a refund record created or an order status returned.
  • Log the outcome of every session - resolved confirmed, resolved assumed, escalated or abandoned, with a fallback flag on every bot reply.
  • Compute each rate with one formula - goal completion rate, escalation rate by reason, fallback rate, and CSAT with its response rate.
  • Write the formula and the unit beside each number - a rate published without its denominator cannot be checked or compared.
  • Add answer quality checks when an LLM generates the answers - hallucination rate and groundedness, scored against a reference for each response.

To show the formulas disagreeing on one set of conversations, this section uses two short Node.js scripts written for this guide:

  • Log generator - writes a sample log of 40 sessions for a fictional support chatbot. It sets the mix of outcomes by hand (13 confirmed, 9 assumed, 8 escalated, 6 abandoned and 4 that never engage) and draws every other field from a fixed random seed, with odds chosen for each outcome.
  • Metrics script - reads that log and prints each metric with its fraction and formula.
  • Sample data - none of the numbers is a measurement of a real chatbot, and none is a benchmark. The gaps in the output illustrate the arithmetic and are not findings.

Each session record carries the fields the formulas need: the platform's outcome label, a back-end goal flag, the labeled and predicted intent, and fallback, latency and token values for every reply. This is one record from the sample log, reformatted for width:

{
  "id": "S016",
  "engaged": true,
  "labeledIntent": "change_address",
  "predictedIntent": "change_address",
  "outcome": "resolved_assumed",
  "escalationReason": null,
  "goalCompleted": false,
  "userTurns": 2,
  "botReplies": [
    { "fallback": false, "latencyMs": 1038, "tokens": 1370 },
    { "fallback": false, "latencyMs": 1759, "tokens": 955 }
  ],
  "csat": null,
  "recontactWithin7d": true
}

Both scripts ran on Node.js 24.21.0 on October 5, 2026, and the second printed this:

Sample log: 40 sessions, 36 engaged, 4 unengaged
Rates are over engaged sessions unless a line says otherwise.

OUTCOME
Goal completion rate              52.8% (19/36)     goal confirmed by the back end / engaged
Resolution rate, confirmed only   36.1% (13/36)     resolved confirmed / engaged
Resolution rate, with assumed     61.1% (22/36)     (confirmed + assumed) / engaged
Containment rate                  77.8% (28/36)     (engaged - escalated) / engaged
Containment rate, all sessions    80.0% (32/40)     (all - escalated) / all, unengaged included
Goal confirmed by the back end, per outcome label:
  resolved confirmed              13 of 13
  resolved assumed                4 of 9
  escalated                       0 of 8
  abandoned                       2 of 6

HANDOFF AND DROP-OFF
Escalation rate                   22.2% (8/36)      escalated / engaged
  user requested                  1
  system unintended               6
  system intended                 1
Abandonment rate                  16.7% (6/36)      abandoned / engaged
Recontact rate, 7 days            27.3% (6/22)      resolved with a repeat contact / resolved

UNDERSTANDING AND EFFORT
Intent recognition accuracy       80.6% (29/36)     first messages matched correctly / labeled
  lowest intent                   57.1% (4/7)       change_address
Fallback rate, per bot reply      14.2% (18/127)    fallback replies / bot replies
Fallback rate, per conversation   30.6% (11/36)     engaged with 1+ fallback / engaged
Average turns to resolution       3.92 (51/13)      user turns in confirmed / confirmed

EXPERIENCE AND COST
CSAT, mean of 1 to 5              3.75 (45/12)      sum of ratings / ratings received
CSAT, satisfied share             58.3% (7/12)      ratings of 4 or 5 / ratings received
CSAT response rate                33.3% (12/36)     ratings received / engaged
Response latency, mean            1,791 ms          over 127 bot replies
Response latency, P50             1,359 ms          nearest-rank percentile
Response latency, P95             6,260 ms          nearest-rank percentile
Tokens per confirmed resolution   11,579            all engaged tokens (150,533) / confirmed
  confirmed sessions only         4,957             their own tokens (64,438) / confirmed

95% WILSON SCORE INTERVALS, 36 ENGAGED SESSIONS
Containment rate                  77.8%             61.9% to 88.3%
Resolution rate, confirmed only   36.1%             22.5% to 52.4%
Goal completion rate              52.8%             37.0% to 68.0%

Every line of that output comes from the same 40 sessions, and the headline changes with the definition:

  • Outcome metrics - containment is 77.8%, resolution with assumed is 61.1%, goal completion is 52.8% and confirmed resolution is 36.1%. Of the 28 contained sessions, the back end confirms the goal in 19.
  • Goal completion by outcome label - all 13 confirmed resolutions have a completed goal, while only 4 of the 9 assumed resolutions do. Two of the 6 abandoned sessions also completed their goal. That happens when a user gets the answer and leaves without confirming.
  • Fallback rate - 14.2% per bot reply and 30.6% per conversation come from the same 18 fallback replies.
  • Unengaged sessions - counting the 4 sessions that never engaged lifts containment from 77.8% to 80.0%, because a session that never asked for anything was never escalated.
  • Intent recognition accuracy - the overall 80.6% includes a change_address intent at 57.1%.
  • CSAT - the 3.75 average rests on 12 ratings. That is a 33.3% response rate.
  • Response latency - the mean of 1,791 ms sits between a P50 of 1,359 ms and a P95 of 6,260 ms and describes neither.
  • Cost per resolved conversation - counting every engaged session's tokens gives 11,579 tokens per confirmed resolution (the script leaves out the single reply in each of the 4 unengaged sessions), and counting only the confirmed sessions' own tokens gives 4,957.

Metrics 13 to 15 are not in the output: they need a reference and a label for each response, which a session log does not hold.

A rate from 36 sessions is an estimate. In Interval Estimation for a Binomial Proportion, Brown, Cai and DasGupta describe the "chaotic coverage properties" of the standard Wald interval and "recommend the Wilson interval or the equal-tailed Jeffreys prior interval for small n". The Wilson interval for containment here runs from 61.9% to 88.3%, so a move of a few points between two weeks of this size is not evidence of a change.

The script computes the interval with this function:

// 95% Wilson score interval for a proportion of k successes in n trials.
function wilson(k, n, z = 1.96) {
  const p = k / n;
  const denom = 1 + (z * z) / n;
  const center = (p + (z * z) / (2 * n)) / denom;
  const half = (z * Math.sqrt((p * (1 - p)) / n + (z * z) / (4 * n * n))) / denom;
  return [center - half, center + half];
}

Setting Targets for Chatbot KPIs

Containment, fallback and CSAT figures from other companies are not comparable with yours, because the definitions behind them differ. Copilot Studio and Intercom Fin count resolutions the user never confirmed, Amazon Lex V2 categorizes a conversation by the result of its final intent, and sessions that never engaged can be counted or left out.

In the sample log above, the same 36 engaged sessions give a resolution rate of 61.1% with assumed resolutions and 36.1% with confirmed ones only. That is a 25-point gap on identical conversations.

The guide to agent performance lists published call-center figures for first contact resolution and CSAT with their sources. Read them as context for your own baseline: that guide cautions that its first contact resolution figure depends on how "solved in one contact" is defined.

Set chatbot KPI targets from your own baseline:

  • Freeze the definitions - write down the event, the unit and the timeout behind each metric, and version them. A changed definition starts a new baseline.
  • Baseline each intent - compute every metric for each intent before the change you want to judge, because an overall figure can look healthy while one intent fails.
  • Attach an interval - print the Wilson interval with every rate and treat a move inside it as no change.
  • Compare like with like - judge a release by its difference from the baseline under the same definitions.
  • Pair the metrics - read each headline number with the metric that exposes its blind spot.
Headline metricRead it withReason
Containment rateGoal completion rateA session that was abandoned still counts as contained.
Resolution rateRecontact rateAn assumed resolution that comes back was not a resolution.
Fallback rateWrong-intent matchesA lower confidence threshold trades fallbacks for wrong answers.
CSATResponse rateThe score describes only the users who answered the survey.
Response latency (P50)P95The median says nothing about the slowest replies.
Hallucination rateRuns per question and who labeledThe rate moves with the number of runs and with the judge.

Chatbot Metrics You Can Test Before Release

Most of the metrics above can be measured before release, each with a different kind of test, and the rest need real users:

  • Simulated conversations - goal completion, escalation and fallback behavior, turn counts and the three answer quality metrics.
  • A labeled utterance set - intent recognition accuracy.
  • A load test - response latency at your expected peak.
  • Production only - containment, resolution, abandonment, recontact, CSAT and cost per resolved conversation.

TestMu AI's Agent Testing runs the first kind of test, simulated conversations, for chat, voice and phone agents. Its testing agents hold multi-turn conversations with your chatbot through the chatbot's own endpoint, as a user would, and score each one on nine quality metrics. Several of them correspond to metrics in this guide:

  • Hallucination Detection - whether the agent invents information not supported by its knowledge base or context (metric 13).
  • Context Awareness - whether the agent retains and correctly uses information from earlier in the conversation (metric 15).
  • Positive User Outcome - whether the interaction is likely to result in the user achieving their goal. It is the closest pre-release reading of goal completion (metric 1).
  • Completeness - whether the response fully addresses the user's question or need. It is the pre-release signal for recontact (metric 6).

The remaining metrics are Bias Detection, Response Quality, Conversation Flow, Tone Consistency and Root-Cause Understanding. A run rolls up into a Green, Yellow or Red verdict, and each metric in it returns:

  • A pass or fail for the scenario - with an evidence excerpt from the conversation that drove it.
  • A score aggregated across the run - one figure for the metric over all scenarios.
  • A confidence level - High, Medium or Low, based on how many scenarios were evaluated.

For groundedness against your own data, Data Validation checks the values a chat agent states, such as an order status or a price, against your system of record through read-only API lookups after each conversation. Each fact gets a Pass, Fail or Cannot verify verdict, and the setup is in the chat agent testing documentation.

A chat run does not cover:

  • Latency under load - a chat run does not model production-scale concurrent usage, and streaming latency is not measured separately. The product's performance testing is for phone agents, so a chatbot needs a separate load test.
  • Intent recognition accuracy - it is not one of the nine chat and voice metrics, so score your labeled utterance set separately.
  • Memory across sessions - continuity between separate sessions is not evaluated.

A Green verdict covers the scenarios, personas and metrics you configured and is no guarantee of zero failures in production.

If your team prefers to write evaluation logic in code, an evaluation library gives you the building blocks, and Agent Testing is for teams that want the scenarios, simulated users and scoring without building them in-house. To turn these checks into a test plan, see how to test a chatbot, which walks through the process step by step.

Note

Note: Beyond the nine standard metrics, TestMu AI Agent Testing lets you define custom pass or fail rules for your own chatbot, and its dashboard shows metric score deltas between test runs. Start free with TestMu AI.

Conclusion

Take the number your team reports most often and write its event, formula, unit and timeout on one line. Then compute goal completion rate for the same sessions from a back-end event and put the two side by side. The gap shows which conversations your current chatbot metrics count as wins without the user getting what they came for.

Before the next prompt, model or knowledge-base change, run the answer quality checks on simulated conversations, so that hallucination rate and context retention have a pre-release value to compare against. TestMu AI's Agent Testing runs them from the dashboard or the CLI, and the guide to testing your first AI agent walks through the first run.

Author

...

Chaitanya Sharma

Blogs: 18

  • Linkedin

Chaitanya Sharma is an AI Product Manager at TestMu AI (formerly LambdaTest), where he builds agentic AI capabilities focused on computer vision and multi-modality, moving testing beyond static script execution toward autonomous, agent-driven workflows. Before TestMu AI he shipped 135+ features at Sprinklr for a no-code community and website builder used by Fortune 500 enterprises including Dell, Samsung, and Polestar. At Policybazaar he led the zero-to-one launch of a digital lending and insurance marketplace embedded in Bahrain's dominant payments app, building a risk-intelligence engine that compressed loan-approval times by 80%. He explored machine learning and NLP through research at the University of Cambridge, and holds a B.Tech from Delhi Technological University.

Reviewer

...

Sai Krishna

Reviewer

  • Linkedin

Sai Krishna is Director of Engineering at TestMu AI (formerly LambdaTest), where he leads agentic AI for quality engineering, building AI agents that autonomously drive mobile and conversational test automation. His current focus is Agent Testing and Model Context Protocol (MCP) support for mobile. He is a core contributor and member of the Appium open-source project and the creator of AppiumTestDistribution and appium-device-farm. With over 14 years of experience including more than 9 years at Thoughtworks as a Principal Consultant, he holds a BSc in Electronics and speaks regularly at TestMu and Appium Conf on Appium, mobile automation, and agentic AI in testing.

Add to Google preferred sources

Summarise with AI

Copied to Clipboard!
...

3000+ Browsers. One Platform.

See exactly how your site performs everywhere.

Try it free
...

Write Tests in Plain English with KaneAI

Create, debug, and evolve tests using natural language.

Try for free

Chatbot Metrics FAQs

Did you find this page helpful?

More Related Learning Hubs

TestMu AI forEnterprise

Get access to solutions built on Enterprise
grade security, privacy, & compliance

  • Advanced access controls
  • Advanced data retention rules
  • Advanced Local Testing
  • Premium Support options
  • Early access to beta features
  • Private Slack Channel
  • Unlimited Manual Accessibility DevTools Tests