Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Learning Hub
- /
- RAG Evaluation Metrics: Formulas and a Worked Example
RAG Evaluation Metrics: Formulas and a Worked Example
RAG evaluation metrics with formulas and a worked example: precision@k, recall@k, MRR, nDCG, faithfulness and answer relevancy, plus why frameworks disagree.
Published on:
OVERVIEW
A support assistant tells a customer that a replacement router ships in two business days, and the policy it retrieved says five. The team's dashboard shows one answer-quality score for the week, and that score cannot say whether the retriever fetched the wrong chunk, the context builder dropped the right one, or the model changed a number that sat in its prompt. RAG evaluation metrics exist to separate those cases.
In a 2024 experience report on engineering RAG systems, researchers at Deakin University's Applied Artificial Intelligence Institute drew on three case studies from separate domains and catalogued 7 failure points, including a relevant document ranked too low to be returned and an answer the model failed to extract from its own context.
Telling those failure points apart takes a score at each stage of the pipeline. This guide works every metric on one request, a warranty question scored against a sample knowledge base on 5 October 2026, so you can see what each metric counts and what moves it. Five more sample questions supply the averages.
Overview
RAG evaluation metrics are scores that measure each stage of a retrieval-augmented generation pipeline. Retrieval metrics check whether the right chunks were fetched and ranked near the top. Generation metrics check whether the answer is supported by those chunks and addresses the question, and reference-based metrics compare the context or the answer with a known correct answer.
Which RAG evaluation metrics should you track?
- Precision@k: Precision@k is the share of the top k retrieved chunks that are relevant. This guide's worked example has 2 relevant chunks in its top 5, which is a precision@5 of 0.400.
- Recall@k: Recall@k is the share of all relevant chunks that appear in the top k. The worked example's top 5 holds 2 of its 3 relevant chunks, which is a recall@5 of 0.667.
- Mean reciprocal rank (MRR): MRR averages, across queries, 1 divided by the rank of the first relevant chunk. A look-alike chunk ranks first in the worked example, so its reciprocal rank is 0.500.
- nDCG@k: nDCG@k compares a ranking's discounted gain with that of an ideal ordering. The worked example scores 0.498 in its original order and 0.765 once the same five chunks are reordered with the relevant ones first.
- Context recall: Context recall is the share of claims in a reference answer that the retrieved context supports. The worked example's context supports 4 of 5 reference claims, which is a hand-worked context recall of 0.80.
- Faithfulness: Faithfulness is the share of the answer's claims that the retrieved context supports. Worked by hand, the example scores 3 of 5, or 0.60, and 0.80 under a rule that penalizes only contradictions.
- Answer relevancy: Answer relevancy measures whether the answer addresses the question, without checking its facts. Worked by hand, the example scores 1.00 while two of its five claims are wrong or unsupported.
- Factual correctness: Factual correctness compares the claims in an answer with a reference answer and reports precision, recall and F1. Worked by hand, the example scores 0.60 on each, with two claims the reference does not make and two reference claims the answer leaves out.
What do RAG evaluation metrics not measure?
RAG evaluation metrics score text: the retrieved chunks and the claims in an answer. They do not show whether an agent that acts on a retrieved policy changed the right record. TestMu AI's Agent Assurance tests that before release: it calls the real agent on staging and grades each criterion on evidence from the run, such as the tool calls made and the records a read-only lookup returns.
What Are RAG Evaluation Metrics?
RAG evaluation metrics are numeric scores, usually between 0 and 1, that each measure one stage of a retrieval-augmented generation pipeline: whether the retriever returned and ranked the right chunks, whether the context that reached the prompt was complete, and whether the answer is supported by that context and addresses the question.
Retrieval metrics score the ranked list of chunks, and generation metrics score the claims in the answer. To plan an evaluation, sort RAG metrics by the evidence each one needs:
- Labeled chunks - precision@k, recall@k, hit rate, MRR and nDCG compare the retrieved chunk ids with the ids a person marked relevant. They run without a model, so the same inputs always return the same score.
- A reference answer - context recall and factual correctness compare the context, or the answer, with an answer written in advance by someone who knows the right one. Judge-based context precision usually asks for one too, to decide which chunks count as relevant.
- The run alone - faithfulness, answer relevancy and context relevance need only the question, the retrieved context and the answer, so they can score live traffic. Each relies on a judge or embedding model, and the score moves when that model changes.
Score both families. In a SIGIR 2024 study of retrieval evaluation in RAG, Alireza Salemi and Hamed Zamani report that scoring a retriever on query-document relevance labels "shows a small correlation with the RAG system's downstream performance", so a strong precision@k is not evidence that the answers are right.
Definitions also differ between frameworks and between versions of one framework, so each formula below names its source and the version read on 5 October 2026. Building the question set, setting thresholds and gating a release on these scores are covered in the guide to RAG testing; this page stays on the metrics themselves.
The Worked Example: A Warranty Question Scored at Each Checkpoint
The example is a support assistant for a fictional maker of network hardware. Its 12-chunk knowledge base, written for this guide, covers warranty terms, return rules, firmware steps and LED meanings, and six questions were labeled by hand with the chunks needed to answer each one.
All of it is sample data, built so that one request shows a ranking fault, a missing chunk and an altered fact together. The ranked lists are part of the sample: the five retrieved chunk ids for each question were written by hand, and no retriever, context builder, reranker or generator model was run.
The request, q1 in the sample, passes through these checkpoints, and each one after the query has its own scores:
- Query - a customer asks, "My R2 router died after 14 months. Is it still under warranty, and how do I get a replacement?"
- Retrieval - a retriever returns chunk ids in ranked order; in this sample, the five ids are a hand-written list. Step 1 scores that list.
- Context assembly - a context builder decides which of those chunks reach the prompt, and in what order. Step 2 scores what survives.
- Generation - a model writes an answer from that context; here the answer was written by hand. Step 3 checks the answer against the context.
- Answer - the customer reads the answer. Step 4 checks it against the question and against the answer a support lead would have written.
The script behind this guide was run on 5 October 2026 with Python 3.14.6 on Windows 11. It uses only the standard library, with no model and no network call, and it prints the sample before any score:
Sample data written for this article: 6 labeled queries, 12 chunks, k=5
q1 My R2 router died after 14 months. Is it still under warranty, and how do I get a replacement?
retrieved: warranty-r1, warranty-r2, returns-30d, claim-steps, doa-exchange
relevant: warranty-r2, claim-steps, exclusions
q2 How long do I have to return an unopened R1?
retrieved: returns-30d, doa-exchange, warranty-r1, led-power, warranty-r2
relevant: returns-30d
q3 Does the warranty cover damage from a power surge?
retrieved: warranty-r2, claim-steps, warranty-r1, returns-30d, exclusions
relevant: exclusions
q4 How do I update the R2 firmware?
retrieved: firmware-check, firmware-install, led-power, warranty-r2, returns-30d
relevant: firmware-check, firmware-install
q5 Can I transfer my warranty to a new owner?
retrieved: warranty-r2, warranty-r1, claim-steps, exclusions, returns-30d
relevant: transfer
q6 What do the LED colors on the R2 mean?
retrieved: led-internet, firmware-check, led-power, warranty-r2, led-wifi
relevant: led-power, led-internet, led-wifi
q1 context in ranked order (R = labeled relevant, . = not relevant)
1 . warranty-r1 The R1 router carries a 12-month limited hardware warranty from the date of purchase.
2 R warranty-r2 The R2 router carries a 24-month limited hardware warranty from the date of purchase.
3 . returns-30d Unopened products can be returned within 30 days of purchase for a full refund.
4 R claim-steps To claim, open a case with the device serial number and proof of purchase. Approved claims ship a replacement within 5 business days.
5 . doa-exchange A device that fails within 30 days of purchase is exchanged immediately as dead on arrival.
- R exclusions The warranty does not cover damage from power surges, liquid, or unauthorized firmware. (not retrieved)Read q1's context from the top: a chunk for the wrong model, warranty-r1, ranks first; the two relevant chunks in the list sit at ranks 2 and 4; and the exclusions chunk that a complete answer needs is not in it. From here on, values with three decimals were printed by this script, and values with two decimals were worked by hand from a documented formula, because the claim-based metrics need a judge model, and none was run.
Small relevant sets like these show up in real data as well. In a 2026 applied study from Orange Research, a human-annotated set of 96 questions over 479 business documents averaged 1.3 reference spans per question, and 15 of the questions had no answer in the documents at all.
Step 1: Score What the Retriever Returned
RAG retrieval metrics compare the ranked chunk ids with the ids labeled relevant, and use nothing else. These are the functions the script applies, which need only Python's math module:
def precision_at_k(retrieved, relevant, k):
return sum(1 for cid in retrieved[:k] if cid in relevant) / k
def recall_at_k(retrieved, relevant, k):
return sum(1 for cid in retrieved[:k] if cid in relevant) / len(relevant)
def hit_at_k(retrieved, relevant, k):
return 1 if any(cid in relevant for cid in retrieved[:k]) else 0
def reciprocal_rank(retrieved, relevant, k):
for rank, cid in enumerate(retrieved[:k], start=1):
if cid in relevant:
return 1 / rank
return 0.0
def ndcg_at_k(retrieved, relevant, k):
# Binary gain, log2(rank + 1) discount. The ideal ranking puts every relevant chunk first.
dcg = sum(1 / math.log2(rank + 1) for rank, cid in enumerate(retrieved[:k], start=1) if cid in relevant)
ideal = sum(1 / math.log2(rank + 1) for rank in range(1, min(k, len(relevant)) + 1))
return dcg / idealFor the six sample queries at k=5, they printed:
query ranks 1-5 relevant P@5 R@5 hit RR nDCG@5
q1 .R.R. 3 0.400 0.667 1 0.500 0.498
q2 R.... 1 0.200 1.000 1 1.000 1.000
q3 ....R 1 0.200 1.000 1 0.200 0.387
q4 RR... 2 0.400 1.000 1 1.000 1.000
q5 ..... 1 0.000 0.000 0 0.000 0.000
q6 R.R.R 3 0.600 1.000 1 1.000 0.885
mean precision@5 0.300 | recall@5 0.778 | hit rate@5 0.833 | MRR 0.617 | nDCG@5 0.628Precision@k and Recall@k
- Precision@k - relevant chunks in the top k, divided by k. For q1, 2 / 5 = 0.400.
- Recall@k - relevant chunks in the top k, divided by all relevant chunks. For q1, 2 / 3 = 0.667.
Introduction to Information Retrieval, the Cambridge University Press textbook by Manning, Raghavan and Schütze, notes that precision at k needs no estimate of how many relevant documents exist, and warns that it "does not average well, since the total number of relevant documents for a query has a strong influence on precision at k". For recall at a cutoff, NIST's trec_eval source uses the same ratio as the script: "relevant retrieved / relevant".
In the sample, query q2 has its only relevant chunk at rank 1 and still scores 0.200, because one relevant chunk in five slots is its ceiling. That is how the same six queries produce a mean precision@5 of 0.300 and a hit rate of 0.833.
- Low precision@k - usually means look-alike chunks crowd the list, as the R1 warranty chunk does for an R2 question, or k is larger than the question needs. Compare the score with its ceiling, relevant chunks divided by k, before treating it as a fault.
- Low recall@k - usually means a needed chunk was not retrieved. The sample leaves the exclusions chunk out of q1's list on purpose: the customer never mentions a cause of failure, so nothing in the question's wording points at that chunk, even though the full answer depends on it.
- Levers - for precision, a metadata filter on the product model, which would exclude the R1 chunk here; for recall, a larger k or different chunk boundaries. Reordering the same k chunks changes neither, and a smaller k is no shortcut, since precision@3 for q1 is 0.333, as Step 2 prints.
- No relevant chunk - recall@k divides by the number of relevant chunks and nDCG by the ideal ranking's gain, and both are zero when no chunk is relevant. Given a question with no relevant chunk, both functions in the script raised a ZeroDivisionError, so keep unanswerable questions in their own set and score them as Step 4 describes.
Hit Rate and Mean Reciprocal Rank
- Hit rate@k - 1 when any relevant chunk is in the top k and 0 otherwise, averaged over queries. Five of the six sample queries hit, which gives 0.833.
- Mean reciprocal rank (MRR) - for each query, 1 divided by the rank of the first relevant chunk, or 0 when none is in the top k; MRR is the mean over queries. The first relevant chunk for q1 is at rank 2, so its reciprocal rank is 0.500, and the sample's MRR is 0.617.
In trec_eval, the file m_recip_rank.c describes reciprocal rank as "most useful for tasks in which there is only one relevant doc, or the user only wants one relevant doc". Its m_success.c file ships hit rate under the name success, a measure of whether "a relevant doc has been retrieved" by a given cutoff.
- Low MRR - the first relevant chunk is buried. The only relevant chunk for q3 sits at rank 5, a reciprocal rank of 0.200, while hit rate and recall both score that retrieval as perfect. Reranking is the lever, since MRR looks only at the first relevant chunk.
- Partial retrieval - hit rate hides it. Query q1 counts as a hit with one of its three relevant chunks missing. Hit rate and recall@5 are the same number for q2, q3 and q5, which have one relevant chunk each, and differ only when a question needs several.
nDCG@k
Normalized discounted cumulative gain (nDCG) adds a gain for each relevant chunk, discounts it by the logarithm of its rank, and divides the total by the same total for an ideal ordering. The script gives every relevant chunk a gain of 1 and discounts by log2(rank + 1). For q1 that is (1/log2(3) + 1/log2(5)) divided by (1/log2(2) + 1/log2(3) + 1/log2(4)), or 1.0616 / 2.1309 = 0.498.
That is the textbook formula with binary labels. Introduction to Information Retrieval writes the gain as 2 to the power of the relevance grade, minus 1, and says NDCG "is designed for situations of non-binary notions of relevance", which makes it the one metric in this step that can use graded labels, such as 2 for an essential chunk and 1 for a helpful one.
- Low nDCG@k - relevant chunks sit low in the list, or some are missing, because the ideal ordering counts every relevant chunk. Query q1 loses on both counts and scores 0.498.
- Levers - reranking, and retrieving the missing chunk. Moving the two relevant chunks that q1 retrieved to ranks 1 and 2 lifts it to 0.765, and only the missing exclusions chunk keeps it below 1.
Step 2: Score the Context That Reaches the Prompt
The prompt does not always hold what the retriever returned. A context builder can drop or reorder chunks to fit a token budget, so score what the model read. The script scores q1 three ways: the five chunks as listed, the first three, which is what a builder with a tighter budget would keep, and the same five in the order a reranker should produce, with the two relevant chunks on top.
q1 under three conditions
view P@k R@k RR nDCG@k AP rank-weighted precision
top 5 retrieved 0.400 0.667 0.500 0.498 0.333 0.500
top 3 kept 0.333 0.333 0.500 0.296 0.167 0.500
top 5 reranked 0.400 0.667 1.000 0.765 0.667 1.000- Trimming - keeping the top three of five cuts recall from 0.667 to 0.333, and nDCG from 0.498 at k=5 to 0.296 at k=3, because the claim-steps chunk at rank 4 no longer reaches the prompt.
- Reranking - precision@5 and recall@5 stay at 0.400 and 0.667, while the rank-aware scores move: reciprocal rank from 0.500 to 1.000 and nDCG@5 from 0.498 to 0.765.
The last column is context precision as two frameworks define it. The Ragas documentation (the stable docs for version 0.4, pages dated 9 December 2025) describes it as the retriever's ability "to rank relevant chunks higher than irrelevant ones": precision at each rank that holds a relevant chunk, summed, then divided by the number of relevant chunks in the top K. The DeepEval documentation (read on 5 October 2026, when the current release was 4.2.8) gives the same formula as contextual precision and says the metric "focuses on evaluating the re-ranker".
In both, a judge model normally decides which chunks are relevant by comparing each one with a reference answer, and Ragas also has a variant that compares with the generated response. The script applies the formula to the hand labels instead, next to textbook average precision, which has the same numerator:
def precision_sum(retrieved, relevant, k):
# Precision at each rank that holds a relevant chunk, summed. Returns (sum, relevant chunks found).
found, total = 0, 0.0
for rank, cid in enumerate(retrieved[:k], start=1):
if cid in relevant:
found += 1
total += found / rank
return total, found
def average_precision(retrieved, relevant, k):
total, _ = precision_sum(retrieved, relevant, k)
return total / len(relevant) # divide by ALL relevant chunks
def rank_weighted_precision(retrieved, relevant, k):
total, found = precision_sum(retrieved, relevant, k)
return total / found if found else 0.0 # divide by the relevant chunks that were RETRIEVEDFor q1 the sum is 1/2 + 2/4 = 1, which gives 0.500 when divided by the two relevant chunks retrieved and 0.333 when divided by all three relevant chunks, as average precision does. After reranking, rank-weighted context precision reads 1.000 while a needed chunk is still missing and three of the five chunks are still irrelevant.
Read it as a score for order. It moves when chunks are reordered, or when an irrelevant chunk ranked above a relevant one is removed, and a missing chunk leaves it unchanged. Average precision, at 0.667 after reranking, still shows the gap.
Context recall asks what the context failed to bring. The Ragas documentation computes it from a reference answer, as the claims in the reference that the retrieved context supports divided by all claims in the reference, and DeepEval's contextual recall does the same with statements from the expected output. The sample's reference answer for q1 has five claims:
- The R2 carries a 24-month limited hardware warranty.
- A failure at 14 months falls inside the warranty period.
- To claim, open a case with the serial number and proof of purchase.
- Approved claims ship a replacement within 5 business days.
- The warranty does not cover damage from power surges, liquid or unauthorized firmware.
The retrieved context supports the first four, and the fifth is in the chunk that was not retrieved, so context recall worked by hand is 4 / 5 = 0.80. The script's recall@5 for the same retrieval is 0.667, because one metric counts claims and the other counts chunks, and the claim count shifts with how the reference answer is written.
A low context recall means a claim the reference needs has no support in the context. Compare it with recall@k for the same question: a chunk that was never retrieved, like the exclusions here, is a retrieval fault, and one that was retrieved and then dropped is a context-assembly fault.
Context relevance asks how much of the context is padding. The 2023 paper that introduced Ragas defines it as the number of context sentences the answer needs, divided by all sentences in the context, and its authors found it "the hardest quality dimension to evaluate". The five chunks retrieved for q1 hold six sentences and three are needed, so the hand-worked value is 3 / 6 = 0.50, and dropping the three irrelevant chunks would raise it to 3 / 3.
Step 3: Check the Answer Against the Context With Faithfulness
Faithfulness is the share of the answer's claims that the retrieved context supports. The Ragas paper computes it as F = |V| / |S|, where S is the set of statements a judge model extracts from the answer and V is the subset it finds supported. The guide to LLM-as-a-judge covers how such judges are built and where they go wrong.
Suppose the model answers q1 with this text, written to carry two faults:
Yes, your R2 is still under warranty. It carries a 24-month limited hardware warranty, so a failure at 14 months falls inside the warranty period. Open a case with your serial number and proof of purchase, and a replacement ships within 2 business days. Shipping is free.Split into claims and checked against the five chunks in q1's context, the answer reads:
- The R2 carries a 24-month limited hardware warranty. Supported by
warranty-r2. - A failure at 14 months falls inside the warranty period. Supported, since it follows from the same chunk.
- To claim, open a case with the serial number and proof of purchase. Supported by
claim-steps. - A replacement ships within 2 business days. Contradicted, because
claim-stepssays 5 business days. - Shipping is free. Absent from the context, which neither supports nor contradicts it.
Three of the five claims can be inferred from the context, so the paper's formula, worked by hand, gives 3 / 5 = 0.60. The current Ragas documentation states the same test, whether each claim "can be inferred from the retrieved context", and the TruLens documentation (read on 5 October 2026) calls that check groundedness. Ragas also documents a separate Response Groundedness metric, which rescales and averages two judge ratings of 0, 1 or 2 instead of counting claims.
DeepEval's documentation states a different rule: "A claim is considered truthful if it does not contradict any facts presented in the retrieval_context." Its judge prompt, in release 4.2.8 and on the main branch of the project's GitHub repository on 5 October 2026, marks a claim the context does not back up as borderline, and the scoring code counts borderline claims as faithful unless you set penalize_ambiguous_claims, which is off by default.
The same documentation page reads the other way in its FAQ, which says a truth absent from the retrieved text "still counts as unfaithful"; that is how the metric scores once the setting is on. Under the stated rule and the default scoring, applied by hand, claim 5 passes and the same answer scores 4 / 5 = 0.80. A real judge may split the answer into different claims or label claim 5 differently, so treat 0.60 and 0.80 as illustrations of two formulas.
- Low faithfulness - the model changed a detail, as in claim 4, or added one from memory, as in claim 5. The source text was in the prompt, so this is a generation fault.
- Levers - more of the needed context, and explicit requirements in the prompt. In tuning experiments on three RAG baselines in the 2024 RAGChecker paper, which prints its scores on a 0 to 100 scale, raising k from 5 to 20 lifted claim recall from 61.5 to 77.6 and faithfulness from 88.1 to 92.2, while noise sensitivity rose from 34.0 to 35.4. Prompts that added explicit requirements for faithfulness, context utilization and lower noise sensitivity then moved faithfulness from 92.2 to 93.6, and noise sensitivity still rose from 35.4 to 38.1.
- Blind spot - faithfulness compares the answer with the retrieved context and with nothing else. An answer copied accurately from a stale chunk scores 1.00 and is wrong, and the TruLens documentation says as much: an application that passes its checks is "hallucination free up to the limit of its knowledge base".
A stale fact needs a check outside the context. For a deployed chat agent, TestMu AI's Agent Testing runs one as Data Validation: the facts the agent states, such as a price or an order status, are compared with your own system of record through read-only API lookups, with a Pass, Fail or Cannot verify verdict for each fact. An answer with no retrieved context behind it is a separate problem, covered in the guide to LLM hallucination detection.
Step 4: Check the Answer Against the Question and a Reference
Answer relevancy asks whether the answer addresses the question, and by design it ignores whether the answer is true. The Ragas paper says its assessment "does not take into account factuality, but penalises cases where the answer is incomplete or where it contains redundant information". Ragas and DeepEval compute it differently:
- Ragas - the documentation generates questions from the answer, three by default, and averages the cosine similarity between each one's embedding and the original question's. It needs a generator model and an embedding model, so it was not computed for q1.
- DeepEval - the documentation extracts the statements in the answer and divides the ones relevant to the question by the total. All five statements in the sample answer concern the warranty or the replacement, so by hand that is 5 / 5 = 1.00.
An answer relevancy of 1.00 therefore sits beside a faithfulness of 0.60 on the same answer. A low value means the answer drifts off the question or answers a neighboring one; factual errors do not lower it.
Factual correctness is the reference-based check. The Ragas documentation counts claims: true positives are answer claims present in the reference, false positives are answer claims absent from it, and false negatives are reference claims missing from the answer. Precision, recall and F1 follow, and F1 is the default.
- True positives - claims 1 to 3 of the answer, which the reference also makes.
- False positives - claims 4 and 5, the 2-day shipping time and the free shipping.
- False negatives - the reference's 5-day shipping time and its exclusions, which the answer never states.
- Scores, worked by hand - precision 3 / 5 = 0.60, recall 3 / 5 = 0.60, F1 0.60.
Low precision points at wrong or extra claims, and low recall at missing ones. For each missing claim, check whether it was in the context: the 5-day figure was, so the generator changed it; the exclusions were not, so retrieval never supplied them.
Questions the corpus cannot answer need their own score, since recall is undefined for them. The RGB benchmark calls the ability negative rejection and measures a rejection rate: how often a model declines to answer when only noisy documents are retrieved. Across the six LLMs its 2023 paper evaluated, ChatGPT among them, the highest rejection rates were 45% in English and 43.33% in Chinese.
RAG Evaluation Metrics With One Name and Several Formulas
The script scored the unchanged list of five chunks for q1 under a second published definition of hit rate, reciprocal rank and nDCG:
q1, top 5 retrieved, same list under other definitions
hit rate, any relevant chunk in the top 5 1.000
hit rate, relevant chunks found / all relevant 0.667
reciprocal rank, first relevant chunk only 0.500
reciprocal rank, averaged over relevant chunks found 0.375
nDCG@5, discount log2(rank + 1) 0.498
nDCG@5, discount log2(rank) from rank 2 0.570Each alternative comes from published code or a paper. LlamaIndex's retrieval metrics source (main branch, read on 5 October 2026) has a granular option for hit rate that divides matches by expected documents, and one for MRR that averages the reciprocal ranks of every relevant document found. The 2002 ACM Transactions on Information Systems paper by Järvelin and Kekäläinen, which trec_eval cites for its nDCG measure, discounts by the logarithm of the rank itself and applies no discount before the rank equal to the logarithm base, while the textbook and trec_eval both use log2(rank + 1).
The context and claim metrics diverge in the same way:
| Name | Source read | What it computes | On q1 |
|---|---|---|---|
| Context precision | Ragas documentation 0.4; DeepEval documentation, as contextual precision | Precision at each relevant rank, summed, divided by the relevant chunks retrieved | 0.500, then 1.000 after reranking |
| Context precision | RAGChecker paper | Relevant chunks divided by k | 0.400 before and after reranking |
| Average precision | Introduction to Information Retrieval; LlamaIndex source | The same sum, divided by all relevant chunks | 0.333, then 0.667 after reranking |
| Context relevance | Ragas paper, 2023 | Sentences needed to answer, divided by all sentences in the context | 0.50 |
| Context relevance | Ragas documentation 0.4 | Two judge ratings on a scale of 0, 1 or 2, rescaled and averaged | Not computed |
| Faithfulness | Ragas paper; Ragas documentation 0.4 | Claims that can be inferred from the context, divided by all claims | 0.60 |
| Faithfulness | DeepEval documentation (stated rule) and 4.2.8 scoring code, default settings | Claims that do not contradict the context, divided by all claims | 0.80 |
| Answer relevancy | Ragas paper; Ragas documentation 0.4 | Mean cosine similarity between the question and questions generated from the answer | Not computed |
| Answer relevancy | DeepEval documentation | Statements relevant to the question, divided by all statements | 1.00 |
The two-decimal values are the hand-worked ones from Steps 2 to 4, and the rows marked not computed need a judge model or an embedding model. Names also drift inside one project: the Ragas documentation page for answer relevancy is titled Response Relevancy, and it imports AnswerRelevancy in the current API and ResponseRelevancy in the legacy one.
A published comparison found the same kind of divergence. The Orange Research study scored one RAG system with metrics from four libraries, and found that DeepEval's contextual precision correlated above 0.5 with recall while Ragas's non-LLM metrics correlated very weakly, which the author says is probably explained by their comparing spans of very different sizes by Levenshtein distance. The author also notes that the study did not set out to compare the metrics, since its results depend strongly on the choices made in its design.
- Record the setup with every score - the library, its version, the judge model and k. Without them, this month's 0.80 and next month's 0.60 may describe the same system.
- Compare like with like - choose one definition per metric for a project, pin the library version in CI, and keep both until you rebaseline. The roundup of RAG evaluation tools compares the libraries themselves.
How to Measure RAG Performance From the Scores Together
To measure RAG performance, read the scores as a set, because each fault leaves a pattern across several of them. The patterns below use the failure-point names from the Deakin University paper; matching scores to those names is this guide's reading of the sample, since the paper does not assign metrics to them.
- Hit rate of 0 - nothing relevant reached the top k, as in q5, whose one relevant chunk,
transfer, is in the corpus and outside the top five. The paper's name for that is Missed the Top Ranked Documents: raise k, then revisit chunking and the embedding model, since no prompt change repairs it. If the document does not exist at all, the name is Missing Content: score whether the assistant declines, and add the content. - Recall falls between the retriever's list and the prompt - q1 drops from 0.667 on five chunks to 0.333 when three are kept. That is Not in Context: the document was retrieved and the consolidation step dropped it.
- Hit rate and recall fine, MRR and nDCG low - the needed chunk is buried, as in q3. Rerank; nothing else has to change.
- Context recall high, faithfulness low - Not Extracted: the answer was present in the context and the model failed to extract it, as with the 2 business days in q1. Work on the prompt, the generator and the noise around the chunk.
- Faithfulness high, factual-correctness recall low - the answer leaves something out. If the missing claim was in the context, the paper calls it Incomplete; if it was not, as with the exclusions in q1, the fault is back in retrieval.
- Faithfulness high, answer relevancy low - Incorrect Specificity: a grounded answer that is not specific enough, or too specific, for what the user needs.
- Every text metric fine, output unusable downstream - Wrong Format: the model ignored an instruction to return a table or a list. None of these metrics measures that, so assert the format in code.
- Faithfulness high, facts wrong in the system of record - the retrieved chunk is stale, a corpus fault that only a check outside the context, such as the one in Step 3, catches.
The same paper's first takeaway is that "validation of a RAG system is only feasible during operation", so keep faithfulness and answer relevancy, which need only the run, scoring a sample of live traffic after release. The roundup of the best AI evaluation tools for production compares platforms that score sampled live traffic. Where an agent decides what to retrieve and may retrieve several times, score each retrieval this way; the guide to agentic RAG covers those architectures.
Testing the Action Behind a Faithful Answer With Agent Assurance
The answer to q1 scored 0.60 on faithfulness, and a rewrite that scored 1.00 would still say nothing about an order record. Once the warranty assistant can create replacement orders through a tool, its answer can repeat the claim-steps chunk exactly while the order is missing, duplicated, or placed for a power-surge claim that the exclusions chunk rules out. A reported action with nothing behind it is the subject of the guide to agent action hallucination.
Test how your agents actually behave across workflows, tools, and actions. TestMu AI's Agent Assurance does that before release: it writes scenarios from the agent's code, a PRD or a knowledge-base folder, then calls the real agent on staging. Agent Assurance is pre-alpha and publicly installable, which means a command or a stored file format can differ from one build to the next.
Set against the steps above, a run is graded on evidence those scores never see:
- Policy criteria - q1's context never held the exclusions chunk, yet an order placed for a power-surge claim would still break the rule. Exploration reads your PRD, policies and knowledge base for intended answers, policies, domain facts and boundaries, and those become acceptance criteria, whatever the agent retrieved on a given run.
- Observed tool calls - a criterion can require that the order tool is never called for a claim the exclusions rule out, and it is graded from the calls the run made. Those calls are compared with the tools the agent itself declares, and a profile that hands back no calls leaves the criterion Unable to Verify.
- Records - the action-side version of Step 3's blind spot is an answer that says the order was placed when the order system holds no such record. A judge can confirm that the order exists and is correct with a read-only query through a tool you approve, which for now has to be exposed on a stdio MCP server.
- Hallucination scenarios - claim 5 in Step 3 was a detail no chunk contained. Hallucination is one of nine categories in the adversarial class, which is generated by default alongside the functional class, and a scenario can list things the agent must never claim.
- Unable to Verify - the comparison table above marks two values as not computed, because no model was run for them. Agent Assurance holds its own verdicts to the same rule: a criterion with nothing to check it against is reported as Unable to Verify and stays outside the pass rate, where it counts neither for nor against the agent.
Most eval and observability tools score what your agent said and recorded. Agent Assurance checks what the run changed, and reports what it could not verify. The two overlap, since Agent Assurance also scores answers with model judges and runs in CI; what separates them is the standard of proof, because an action the agent only reports is not treated as evidence that it happened.
A run reports a verdict per criterion, a pass rate over the scenarios that reached a verdict, and verification coverage, and none of those is a retrieval score. The Agent Assurance overview states the limit: documentation can show what the agent should know, and it does not prove which documents the deployed agent retrieved. Keep Steps 1 and 2 for that question.
Agent Assurance installs as Rook CLI, a command-line tool, and the npm route needs Node.js 22 or newer:
npm install -g @testmuai/rookIt runs on macOS and Linux, and 64-bit Windows through npm or WSL. The Rook CLI install guide covers the Homebrew and shell-installer routes.
To drive it from Claude Code, install the rook skill:
npx @testmuai/rook-skill@latest install --agent claude-codeThen type a request in the Claude Code chat. One built on q1 might read:
/rook Use knowledge/ as the approved source for warranty rules and cover the replacement-order flow for an R2 router that failed at 14 months. Propose a small functional and adversarial suite, including one case where the customer mentions a power surge. Give that case a criterion that the order tool is never called, and give the others a criterion that confirms the replacement order through the read-only order lookup. Wait for my approval before generating anything, and do not invoke the target yet.Keep the profile on staging, because a replacement order created during a run is a real order and nothing rolls it back. A finished run exits 0 with failing scenarios as well as passing ones, so a CI job has to read the verdicts in rook report --json.
Note: If your retrieval-backed agent also creates orders or changes records, test what each run changed before it ships. Get started with Agent Assurance
Conclusion
Run the retrieval functions from Step 1 on what your retriever returns today, for a set of your own questions labeled with the chunk ids that answer them. That gives you the first half of your RAG evaluation metrics, precision@k, recall@k, hit rate, MRR and nDCG, with no judge model and no API key.
Then add faithfulness as the first claim-based metric, and write the library, version, judge model and k beside every score. If the assistant also acts on what it retrieves, test those actions before release with TestMu AI's Agent Assurance; the Agent Assurance quickstart walks through a first suite against a sample agent.
Author
Anubhav Singhmaar is an AI Product Manager at TestMu AI driving Kane CLI, the command-line tool that brings browser automation to the terminal, turning natural-language flows into runs in a real Chrome browser that return pass or fail with shareable proof. He owns the roadmap and prioritization and works with engineering to ship developer-facing features. Before TestMu AI, he spent over four years at Sprinklr owning enterprise voice AI across APAC and EMEA. A mechanical engineer turned product manager, he grounds guidance in real QA workflows.
Reviewer
Chaitanya Sharma is an AI Product Manager at TestMu AI (formerly LambdaTest), where he builds agentic AI capabilities focused on computer vision and multi-modality, moving testing beyond static script execution toward autonomous, agent-driven workflows. Before TestMu AI he shipped 135+ features at Sprinklr for a no-code community and website builder used by Fortune 500 enterprises including Dell, Samsung, and Polestar. At Policybazaar he led the zero-to-one launch of a digital lending and insurance marketplace embedded in Bahrain's dominant payments app, building a risk-intelligence engine that compressed loan-approval times by 80%. He explored machine learning and NLP through research at the University of Cambridge, and holds a B.Tech from Delhi Technological University.
RAG Evaluation Metrics FAQs
Did you find this page helpful?
More Related Learning Hubs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests




