Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Learning Hub
- /
- Chatbot Evaluation: What to Score and Who Should Score It
Chatbot Evaluation: What to Score and Who Should Score It
Chatbot evaluation scores conversations against set criteria. Learn what to score, who scores, a rubric, how to check a judge model and how many chats you need.
Published on:
OVERVIEW
"Nearly one in five consumers who have used AI for customer service saw no benefits from the experience," says the Qualtrics announcement of its 2026 Consumer Experience Trends Report, published on 7 October 2025. It calls that "a failure rate almost four times higher than for AI use in general." Qualtrics, which sells experience management software, surveyed more than 20,000 consumers across 14 countries in the third quarter of 2025.
A team that releases a chatbot without an agreed way to score it ends up arguing from transcripts. One person reads ten transcripts and approves the release, another reads ten different transcripts and objects, and neither can show which of them is right. Chatbot evaluation replaces that argument with a procedure.
The sections below are that procedure, for whoever has to say, with evidence, whether a chatbot is good enough. They follow one example throughout: a retail support chatbot that answers order status, refund, delivery and account questions. Every step works with a spreadsheet.
Overview
Chatbot evaluation is the practice of scoring a chatbot's conversations against written criteria to decide whether the bot is good enough to release. You choose the unit to score, the scorer, a rubric with described levels and a set of conversations large enough to trust, and you write the release rule before the run.
Decisions that make up a chatbot evaluation
- Unit of scoring: A chatbot evaluation can score a single reply, a whole conversation or the outcome for the customer. Most rubric criteria belong to the whole conversation, because a conversation with five good replies and one wrong refund amount has still failed.
- Scorer: Rules check facts that have one right answer, and human reviewers set the standard. A judge model applies that standard at volume, and the user reports whether the conversation worked. Each scorer sees something the others miss.
- Rubric: A chatbot evaluation rubric lists five or six criteria and describes what a pass, a partial and a fail look like in a transcript for each one, so two reviewers reach the same verdict on the same conversation.
- Judge model check: Before a judge model's scores are trusted, two people label a sample, and the judge's verdicts are compared with theirs. The share of failures the judge caught is reported separately from the share of passes it agreed with.
- Sample size: A pass rate of 80% measured on 100 conversations has a margin of error of 7.8 points, and on 400 conversations 3.9 points. Small changes between two chatbot versions need hundreds of conversations to show.
- Release rule: Each rubric criterion is assigned a consequence before the run. A failure either blocks the release, must be fixed by the next release, or is watched over time.
Can I evaluate a chatbot without writing code?
Yes. The procedure needs a sample of conversations, a written rubric and two reviewers, and a spreadsheet is enough to record the verdicts and count agreement. To automate the conversations, a platform such as TestMu AI Agent Testing holds multi-turn conversations with a chatbot through the chatbot's API and shows the scores in a dashboard, with no change to the chatbot's code.
What Is Chatbot Evaluation, and How Is It Different From Chatbot Testing?
Chatbot evaluation scores a set of a chatbot's conversations against written criteria, on enough conversations to trust the result, so that a team can say with evidence whether the bot is good enough to release and whether a change made it better or worse.
Testing and evaluation ask different questions of the same chatbot:
- Chatbot testing - checks expected behavior case by case. A test case has an input and an expected result, such as "a customer who asks for a person is handed off", and it passes or fails. The guides to chatbot testing and how to test a chatbot cover the test types, the test cases and the automation.
- Chatbot evaluation - scores quality across many conversations. The same reply can be correct and still incomplete, off-policy or badly worded, so each conversation is scored against several criteria and the scores are read as rates over the whole set.
A team needs both. Tests catch a broken handoff or a missing integration before any scoring starts. Evaluation tells you how often the retail support chatbot gives a correct, complete refund answer when the replies are generated and no two are worded alike.
What Should a Chatbot Evaluation Score: a Reply, a Whole Conversation or the Outcome?
Score all three units, because each answers a different question. A single reply shows whether one answer was right, a whole conversation shows whether the bot kept track and finished the job, and the outcome shows whether the customer's problem was solved. Most rubric criteria belong at the conversation level.
| Unit scored | Example criterion for the retail bot | What the score shows | What the score hides | Who can score it |
|---|---|---|---|---|
| A single reply | The refund window stated in the reply matches the written policy | The exact turn where an answer went wrong | Whether the bot remembered the order number or contradicted itself two turns later | A rule, a reviewer or a judge model |
| A whole conversation | The bot used the order number given in turn one and handed off when the customer asked for a person | Memory, consistency, policy and handoff across turns | Whether the refund was issued after the chat ended | A reviewer or a judge model that reads the full transcript |
| The outcome | The customer did not contact support again about the same order | Whether the customer's problem was solved | Why it was not solved, and which reply caused it | Your logs and order system, or the customer |
An average over replies can hide a failed conversation. Take a six-turn refund conversation in which five replies are correct and one states the wrong refund amount: five of six replies pass, which is 83%, and the customer was still told the wrong amount.
The research benchmark MT-Bench-101 (ACL 2024) builds that case into its scoring. Its authors built "4208 turns across 1388 multi-turn dialogues in 13 distinct tasks", had a model score every turn, and "use the lowest round score as the total score for the dialogue". The benchmark compares general-purpose language models and says nothing about support bots, but the rule transfers: for any criterion where one wrong reply does the damage, the conversation gets the score of its worst turn.
Outcome rates such as resolution, containment and recontact have their own formulas, which are in the guide to chatbot metrics. If the chatbot also takes actions through tools, such as issuing the refund itself, score those actions with AI agent evaluation metrics.
Who Should Score Chatbot Conversations: Rules, Reviewers, a Judge Model or the User?
Use each scorer for what it can see. Rules check facts and formats that have one right answer, human reviewers set the standard, a judge model applies that standard at volume once it has been checked against the reviewers, and the user tells you whether the conversation worked for them.
| Scorer | Good at | Fails at | Relative cost | When to use |
|---|---|---|---|---|
| A rule in code | Values with one right answer: an order status, a refund window, a required disclosure, a handoff that did or did not happen | Anything that depends on wording, such as completeness or tone | Lowest for each conversation, after the rule is written | On every conversation, first |
| A human reviewer | Judging whether the question asked was answered, and writing the standard the other scorers follow | Volume and consistency: two reviewers disagree, and one reviewer tires | Highest for each conversation | On the sample that defines the standard, and on every verdict the others cannot settle |
| A judge model | Applying a written rubric to thousands of transcripts in the same way each time | Facts it has no reference for, and failures it was never told to look for | Low for each conversation, plus the cost of checking it against reviewers | At volume, after the check in this guide |
| The user | Saying whether the conversation worked for them, through a rating or by coming back with the same problem | Coverage and accuracy: few users rate, and a user cannot tell that a confident answer was wrong | None for each conversation, but the signal arrives after release | In production, as a check on the other three |
Word-overlap scores such as BLEU and ROUGE are not in the table. They count the words a reply shares with a reference reply, and a correct refund answer can share almost no words with the reference answer.
In the EMNLP 2016 paper How NOT To Evaluate Your Dialogue System, Liu and colleagues report that "these metrics correlate very weakly with human judgements in the non-technical Twitter domain, and not at all in the technical Ubuntu domain." Overlap cannot stand in for a reader.
The user is a scorer that outside readers cannot fully replace. In a 2024 preprint, Svikhnushina and Pu compared the ratings people gave after talking to a chatbot with the ratings outside readers gave to the same transcripts.
In the authors' dataset of 1920 dialogues with four chatbots, the two sets of scores correlated at 0.32 for a single conversation and at 0.97 (Pearson) when averaged for each chatbot. The authors conclude that "third-party assessments do not effectively mirror first-party user experiences".
Those chatbots were open-domain empathetic bots and the paper has not been peer reviewed, so treat the numbers as a direction. Averaged over many conversations, transcript review moves with what users report, and it is a weak guide to how one particular customer felt.
With only four chatbots in the study, the averaged figure cannot show that outside readers would rank two versions of your bot the way your users would. Keep a user signal in production for both reasons.
How Do You Build a Chatbot Evaluation Set From Real Conversations?
Start from logged conversations, sample them so every topic and every kind of ending is represented, keep the failures, strip personal data, add written cases for rare and risky situations that the logs do not contain, and freeze the result as a numbered version that every later run uses.
- Sample by topic - a random sample of the retail bot's logs would be mostly order status questions. Take a fixed number from each topic (order status, refund, delivery, account) so that a refund problem is not buried under easy lookups.
- Sample by ending - inside each topic, take conversations that ended in a resolution, in a handoff, in the customer leaving mid-conversation and in a low rating. Conversations that ended badly hold most of the failures you need to see.
- Keep the known failures - every escalated complaint and every wrong answer a support agent had to correct goes into the set, with a note on what the right reply was.
- Remove personal data - replace names, addresses, emails, phone numbers and order numbers with made-up values of the same format before any reviewer or model reads the transcripts.
- Add written cases for what the logs lack - a chargeback threat, a request the policy forbids, an order that does not exist. Write the customer's side yourself or have a model draft it, and label these cases as synthetic so their scores can be read separately.
- Include long conversations - short exchanges do not test memory. On the research benchmark LongMemEval (ICLR 2025), built from 500 questions about long chat histories, the authors report "commercial chat assistants and long-context LLMs showing a 30% accuracy drop on memorizing information across sustained interactions."
- Record the source of record for each case - the policy passage, the order record or the account state that the right answer depends on. The factual criterion is scored against it.
- Freeze and version the set - name it, date it and stop editing it. Two chatbot versions can only be compared on the same conversations, and new cases go into the next version of the set.
A logged conversation records what the old chatbot replied. To score a new version, replay the customer's turns against it, by hand or with a simulated user that follows the same goal. How evaluation datasets are structured and stored is covered in the guide to LLM evaluation, and the general vocabulary of graders and datasets in the guide to AI evals.
What Does a Chatbot Evaluation Rubric Look Like?
A chatbot evaluation rubric is a table with one row for each criterion and one column for each level, where every cell describes what a reviewer would see in the transcript. Five or six criteria and the levels pass, partial and fail are enough for a support bot.
Each conversation of the retail support chatbot gets one verdict for each row:
| Criterion | Pass | Partial | Fail |
|---|---|---|---|
| Answered the question asked | Every question the customer asked gets a direct answer or a stated reason why the bot cannot answer | The main question is answered and a second question in the same message is ignored | The reply answers a different question, or returns generic help text |
| Factually correct against the source of record | Every stated order status, date, amount and policy term matches the order system or the written policy | Facts are correct, and one detail the customer needed is missing, such as the refund method | Any stated fact contradicts the source of record, or no source supports it |
| Used earlier turns | The bot uses the order number, item and problem the customer already gave | The bot asks again for one detail it was already given | The bot contradicts an earlier reply or loses track of which order is discussed |
| Stayed in scope and policy | The bot offers only what the policy allows and declines requests outside its job | The bot stays within policy and adds unrequested advice outside its job | The bot promises an exception, a discount or a deadline the policy does not contain |
| Handed off when it should | The bot hands off when the customer asks for a person or the case meets a written handoff condition, and passes the details along | The handoff happens after the customer has to ask twice, or without the details | The bot keeps the customer in the chat when a handoff condition is met, or hands off a question it could answer |
| Tone | Plain, polite wording, with an acknowledgment when the customer reports a problem | Correct wording that is stiff, or an apology repeated in every turn | Blaming, dismissive or sarcastic wording |
Factual accuracy is one row here and a subject of its own: the guide to chatbot hallucination has the question types that expose made-up answers. For a chatbot that answers from retrieved documents, the retrieval step can be scored separately with RAG evaluation metrics.
Three described levels work better than a 1 to 5 scale. A reviewer can point to the sentence that makes a conversation a fail. Few reviewers can say what separates a 3 from a 4, and a release rule needs a verdict it can count. The published guidance leans the same way:
- Binary labels for specific behaviors - in ABC-Eval (ACL 2023), "Annotators provide binary labels on the turn-level indicating the presence or absence of a particular chat characteristic." The authors found the method "more suitable than alternative Likert-style or comparative approaches for dimensional evaluation" of the open-domain chatbots they studied.
- Pass or fail for a judge model - OpenAI's evaluation best practices recommend, for a model used as a judge: "Use pairwise comparison or pass/fail for more reliability".
- A partial level for triage - partial keeps a conversation that needs a wording fix apart from one that misled a customer. Count partial as a pass for criteria you watch and as a fail for criteria that block a release.
Do not expect 100%. The paper that introduced Google's Meena chatbot, Towards a Human-like Open-Domain Chatbot (2020), used a two-question rubric: "We ask human judges to label every model response on these two criteria", sensibleness and specificity.
In the Meena paper, the full version of the chatbot scored 79%, which the authors note "is still below the 86% SSA achieved by an average human". People were rated on open-ended chat there, a different job from support, and they still lost 14 points on a two-question rubric.
How Do You Check a Judge Model Against Human Scores?
Have two people label the same sample, measure how often they agree, then let the judge model score those conversations and compare its verdicts with the agreed human verdict. Report how many passes it agreed with and how many failures it caught as two separate numbers.
- Two reviewers label an overlap sample on their own - each applies the rubric to the same conversations without seeing the other's verdicts.
- Measure the reviewers' agreement first - count the share of identical verdicts and compute Cohen's kappa, which corrects that share for the agreement two raters would reach by chance. If the reviewers disagree often, the rubric is unclear, and no judge can be checked against it yet.
- Settle the differences - the reviewers discuss each disagreement, record one agreed verdict and reword the rubric cell that caused it.
- Run the judge on the same conversations - give it the rubric and, for the factual criterion, the source of record for each case.
- Compare the judge with the agreed verdict in a two-by-two table - report the passes it agreed with and the failures it caught separately, and list the misses by topic.
- Fix the judge's instructions where it misses - then rerun it on the sample, and keep some labelled conversations aside that the instructions were never tuned on.
- Route what the judge cannot settle to a person - verdicts it marks as uncertain, and every verdict in a topic where it missed failures, go to a reviewer.
The threshold for step two comes from health research. In Interrater reliability: the kappa statistic (Biochemia Medica, 2012), McHugh writes that "any kappa below 0.60 indicates inadequate agreement among the raters and little confidence should be placed in the study results", and that "many texts recommend 80% agreement as the minimum acceptable interrater agreement." A chatbot release is not a clinical study, so treat 0.60 as a floor to reach before moving on.
Reporting two rates in step five is the method Hamel Husain describes in his guide to judge models: "Raw agreement can be misleading when classes are imbalanced. Treat the human labels as ground truth and report the judge's True Positive Rate and True Negative Rate separately." It is one practitioner's method, and the sample below shows why it matters.
A worked check on sample data. A script was written for this guide and run on 8 October 2026 with Node.js v25.5.0. Its input is sample data: 40 imaginary conversations of the retail support chatbot in four topics, each with a pass or fail verdict from two reviewers, the verdict they agreed on and a judge's verdict.
All verdicts were typed by hand, the judge's included, and no model was run. The script printed this:
SAMPLE DATA: 40 support conversations, labelled by hand
TWO HUMAN REVIEWERS
Same verdict 37/40 92.5%
Cohen's kappa 0.79
JUDGE MODEL AGAINST THE AGREED HUMAN VERDICT
judge pass judge fail
human pass (30) 28 2
human fail (10) 4 6
Same verdict 34/40 85.0%
Cohen's kappa 0.57
Passes the judge agreed with 28/30 93.3%
Failures the judge caught 6/10 60.0%
PASS RATE EACH SCORER WOULD REPORT
Judge model 32/40 80.0%
Agreed human verdict 30/40 75.0%
FAILURES THE JUDGE PASSED, BY TOPIC
order_status 0 of 2 failures
refund 3 of 5 failures
delivery 1 of 2 failures
account 0 of 1 failures- Agreement and failures caught - the judge gave the same verdict as the reviewers on 34 of 40 conversations, which is 85.0%. It caught 6 of the 10 failures, which is 60.0%. Most conversations pass, so a judge that passes too easily still agrees often.
- Kappa - the two reviewers agreed on 37 of 40 conversations with a kappa of 0.79. The judge's 85.0% agreement is a kappa of 0.57, under the 0.60 floor quoted above. The reviewers' agreement with each other is the ceiling to compare a judge against.
- Reported pass rate - the judge would report 80.0% and the reviewers 75.0% on the same conversations. The judge passed 4 conversations the reviewers failed and failed 2 they passed, so its pass rate comes out five points higher.
- Misses by topic - 3 of the 4 failures the judge passed are refund conversations. That points to the judge's refund instructions as the thing to fix, and to refund conversations as the topic a person reviews until it is fixed.
The verdicts are sample data that measure no real chatbot, judge model or product, and 60.0% is not a typical figure for judge models. Forty conversations is also a small sample, so the five-point gap describes these 40 conversations and is not a measured difference between scorers in general. The same check works criterion by criterion on a full rubric.
Give the judge a reference for facts. In the preprint No Free Labels (last revised on 28 September 2026), financial professionals wrote 160 business and finance questions, and experts graded 1,200 model responses to those questions and to a corrected subset of the MT-Bench benchmark. The authors found that "when not provided with a correct reference, judges show high agreement with human experts only on questions the judges were able to correctly answer themselves." A judge cannot know your refund policy unless the policy is in front of it.
Do not let the chatbot's own model be the judge. In Self-Preference of AI Judges (30 September 2026), Arena, which runs a model comparison platform, describes how "we gave 12 models 1,460 real Text Arena battles" and asked each to pick the better answer. "On average, a model picked its own answer 58% of the time, while people picked that same answer 34% of the time." This is a vendor's blog post about answers to general prompts from the platform's users, so the sizes may differ for support conversations.
How to write the judge's instructions, and the biases to design around, are covered in the guide to LLM-as-a-judge.
How Many Conversations Does a Chatbot Evaluation Need?
A chatbot evaluation needs hundreds of conversations to detect a change of a few points, and 20 conversations can only show a large one. A pass rate near 80% measured on 100 conversations carries a margin of error of 7.8 points, and on 400 conversations the margin is still 3.9 points.
The margin of error of a measured pass rate, at the 95% level, is 1.96 x sqrt(p x (1 - p) / n), where p is the pass rate and n is the number of conversations. The figures below are that formula's results, computed by a second script written for this guide. The formula is the normal approximation, which is rough for small samples and for rates close to 0% or 100%.
| Conversations scored | Margin of error at an 80% pass rate | Margin of error at a 90% pass rate |
|---|---|---|
| 20 | +/- 17.5 points | +/- 13.1 points |
| 50 | +/- 11.1 points | +/- 8.3 points |
| 100 | +/- 7.8 points | +/- 5.9 points |
| 200 | +/- 5.5 points | +/- 4.2 points |
| 400 | +/- 3.9 points | +/- 2.9 points |
| 1,000 | +/- 2.5 points | +/- 1.9 points |
- A sample of 20 - at 20 conversations an 80% pass rate could be anywhere within 17.5 points either way. A move from 80% to 83% is far inside that range. Even 400 conversations leave a margin of 3.9 points.
- Size each topic, as well as the total - the margin for one topic depends on the number of conversations in that topic. The topics whose failures block a release need the most conversations.
- Compare two versions on the same conversations - run the old and the new chatbot on the same frozen set and list the conversations whose verdict changed. Ten conversations that went from pass to fail are a finding you can read one by one, whatever the two pass rates say.
- Repeat each conversation - a generated reply is one draw from many possible replies. For the number of repeats, see testing non-deterministic AI outputs.
The table sizes a pass rate. Finding the kinds of failure in the first place takes fewer conversations: Hamel Husain's rule of thumb in the guide cited above is to "start with around 30 examples and keep going until I do not see any new failure modes".
Should You Evaluate a Chatbot Before Release, in Production, or Both?
Both, with the same rubric. Before release, a frozen evaluation set tells you whether a change is safe to ship. In production, a regular sample of live conversations tells you whether the set still resembles what customers ask, and every failure found there is added to the set.
| Part of the evaluation | Before release | In production |
|---|---|---|
| Where the conversations come from | The frozen evaluation set, replayed against the new version | A sample of finished live conversations, drawn by topic and by ending |
| Unit scored | Each reply and the whole conversation | The whole conversation and the outcome |
| Scorer | Rules and the checked judge model, with reviewers on what the judge cannot settle | The same rules and judge, plus user ratings and a weekly batch read by reviewers |
| How often | On every change, before it ships | Continuously or on a fixed schedule |
| What triggers a rerun | A knowledge base edit, a prompt or model change, or a new intent | A drop in a criterion's pass rate, or a topic the set does not cover |
In production the unit is the finished conversation, and the scoring has to wait for it to finish. LangSmith's documentation describes that design for its own evaluators, which "evaluate entire conversations between a human and an agent, not just individual exchanges" after an idle period marks the conversation as complete: "multi-turn evaluators run once per completed thread, not once per trace."
A failure found in production goes back into the set like this:
- A reviewer confirms the failure against the rubric and records which criterion failed.
- The conversation is stripped of personal data and added to the next version of the evaluation set, with the source of record for the right answer.
- If the judge passed that conversation, it also joins the labelled sample used to check the judge.
How Do Chatbot Evaluation Scores Become a Release Decision?
Write the rule before the run: for each rubric criterion, decide whether a failure blocks the release, must be fixed by the next release, or is only watched. Then compare the scores with those rules and record who approved the result, so the decision does not depend on one averaged number.
This is the rule table for the retail support chatbot, filled in with an illustrative run of 200 conversations. The thresholds and the results are made-up example values that show how the table is read:
| Rubric criterion | Release rule | Example threshold | Example result and decision |
|---|---|---|---|
| Factually correct against the source of record | Blocks the release | No fail on an amount, a date or a policy term | 2 fails, both on refund amounts. Blocked |
| Stayed in scope and policy | Blocks the release | No fail | 0 fails. Clear |
| Handed off when it should | Blocks the release | No fail where the customer asked for a person | 0 fails. Clear |
| Answered the question asked | Fix by the next release | At least 85% pass | 81% pass. Fix scheduled |
| Used earlier turns | Fix by the next release | At least 85% pass | 90% pass. Clear |
| Tone | Watch | No threshold, trend only | 94% pass. Logged |
- Blocking criteria - one customer told the wrong refund amount is a failure whatever the percentage, so the threshold is a number of fails, and each fail is read by a person before the decision.
- Rate criteria - 81% against an 85% threshold on 200 conversations is a gap of 4 points, inside the 5.5 point margin of error listed earlier for 200 conversations at an 80% pass rate. A gap that size cannot be told apart from sampling noise, which is why the rule written before the run for rate criteria is a fix by the next release, and a miss does not block the release.
- The release record - the set version, the chatbot version, the scores, the fails that were read and the name of whoever approved the release are stored together. The next evaluation starts from that record.
In this example the release is blocked by the two refund amounts. The fix goes in, the frozen set is rerun, and the two failed conversations stay in the set for every later version.
Which Chatbot Evaluations Can You Run on TestMu AI Agent Testing?
You can run a scenario-based evaluation of a chat agent, score every conversation on the quality metrics you select, set your own threshold for each metric, check the facts the chatbot states against your own system, and rerun the same suite after every change.
These evaluations run on Agent Testing, TestMu AI's platform for chat, voice, phone (inbound and outbound), video and image agents. For a chat agent you give it the chatbot's HTTP endpoint, a prompt that describes the intended behavior, and your requirement documents. It then holds multi-turn conversations with the chatbot and scores each one, with no change to the chatbot's code.
Step by step on the platform:
- The evaluation set - you write a scenario by hand with a title, a description and the expected behavior, assign it a persona such as a frustrated customer or a first-time user, and link test data to it. The platform also generates 60 to 100+ scenarios from your prompt and documents, spread over happy paths, edge cases, adversarial inputs, personas and compliance checks.
- The criteria - a chat conversation is scored on nine quality metrics, including hallucination detection, bias detection, completeness, context awareness, response quality and conversation flow. You choose which metrics a run scores, or run all of them.
- Your own rubric rows - a validation criterion is a pass or fail rule you write for a scenario. The documentation's example is "Agent must mention return policy". The result for each criterion is Pass, Fail or Unable to Verify, and it comes with the evidence and a confidence level.
- The threshold - you set a minimum score for each metric and whether higher or lower is better, and save the set as a named threshold configuration. The documentation's examples are named "Strict" and "Default", and threshold profiles can differ between development, staging and production.
- Facts against the source of record - for the Chat agent type, Data Validation maps each fact the chatbot may state, such as an order status or a delivery date, to the matching value in an API of yours that it only reads, and fetches that value at evaluation time. Each fact is reported as Pass, Fail or Cannot verify (the Data Validation label for a fact it could not check), with the chatbot's statement beside your API's value. Matching is by meaning: in the documentation's example, "shipped" matches a status of SHIPPED_IN_TRANSIT. The setup is in the chat agent testing documentation.
- What one result contains - the overall score, a score and a pass or fail badge for each metric, a written analysis, the full multi-turn transcript, identified strengths, areas for improvement, recommendations and the validation criteria results.
- The release decision - a run rolls up into a Green, Yellow or Red go-live verdict with an overall score, a confidence level, scenario coverage and a risk assessment. The verdict reports pass and fail rates for each behavioral category, so a Yellow or Red result points at the category to fix. The confidence level rises with the number of evaluations behind the score, as the documentation on quality dimensions and go-live readiness explains.
- The rerun - scheduled runs repeat a suite on a cron schedule, with timezone support, pause and resume controls and a history of runs, and the dashboard shows how each metric's score changed between runs. The
agent-testing-clitool starts a chat evaluation from a terminal or a CI job. The run is asynchronous, and the scores appear in the dashboard.
The platform leaves these parts of the guide to you:
- Scores from live customer conversations - a run scores simulated conversations. Outcome measures from real users, such as resolution and customer ratings, come from your own analytics.
- Facts that no API returns, and actions - Data Validation covers values a read-only API can return as JSON and that the chatbot states in the conversation. It is not built to confirm an action the chatbot took, and a claim about a policy with no API value falls to the metrics and to the validation criteria you write.
- Agreement with your own reviewers - have two reviewers label a sample of scored transcripts against your rubric and compare their verdicts with the platform's scores, counting agreement on passes and on failures separately. A chatbot in a specialized domain may also need more of its own validation criteria and hand-written scenarios.
To begin a chatbot evaluation of your own, write the six rubric rows for your chatbot and pick the 20 logged conversations that worried you most. The docs guide to testing your first AI agent walks through setting up and running simulated conversations on TestMu AI.
Note: Before a full run, the Playground in TestMu AI Agent Testing lets you chat with your configured chat agent turn by turn and test the connection to its endpoint. See the chatbot testing platform.
Author
Chaitanya Sharma is an AI Product Manager at TestMu AI (formerly LambdaTest), where he builds agentic AI capabilities focused on computer vision and multi-modality, moving testing beyond static script execution toward autonomous, agent-driven workflows. Before TestMu AI he shipped 135+ features at Sprinklr for a no-code community and website builder used by Fortune 500 enterprises including Dell, Samsung, and Polestar. At Policybazaar he led the zero-to-one launch of a digital lending and insurance marketplace embedded in Bahrain's dominant payments app, building a risk-intelligence engine that compressed loan-approval times by 80%. He explored machine learning and NLP through research at the University of Cambridge, and holds a B.Tech from Delhi Technological University.
Reviewer
Sai Krishna is Director of Engineering at TestMu AI (formerly LambdaTest), where he leads agentic AI for quality engineering, building AI agents that autonomously drive mobile and conversational test automation. His current focus is Agent Testing and Model Context Protocol (MCP) support for mobile. He is a core contributor and member of the Appium open-source project and the creator of AppiumTestDistribution and appium-device-farm. With over 14 years of experience including more than 9 years at Thoughtworks as a Principal Consultant, he holds a BSc in Electronics and speaks regularly at TestMu and Appium Conf on Appium, mobile automation, and agentic AI in testing.
Chatbot Evaluation FAQs
Did you find this page helpful?
More Related Learning Hubs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




