Attribute the layer before you change the pipeline
When a RAG assistant returns a wrong answer, the argument that follows is usually about whether to swap the embedding model or rewrite the answer prompt. The current documentation answers that with a two-stage split rather than one aggregate quality score. Evidently's guide frames debugging as explicitly asking whether the system failed to retrieve the right documents or whether the model hallucinated even though the context was correct, and groups its metrics left to right from retrieval to generation.[1] The survey 'Evaluation of Retrieval-Augmented Generation' argues that assessment is hard precisely because of RAG's hybrid structure and its reliance on dynamic knowledge sources, and proposes a unified process that compares quantifiable metrics of the retrieval and generation components separately.[3]
LangSmith's tutorial turns that split into mechanics you can implement. It defines four evaluators, each comparing exactly two artifacts drawn from the question, the response, the retrieved documents, and a reference answer: correctness, relevance, groundedness, and retrieval relevance.[2] The artifact-pair framing is what makes a score attributable to a layer. Retrieval relevance never looks at the answer, so it cannot be contaminated by generation quality, and groundedness never looks at the question, so it cannot be dragged down by a retrieval miss.[2]
The practical consequence is a logging requirement, not a dashboard requirement. If your traces store only the question and the final answer, no metric you add later can separate the layers, because the comparison it needs is missing. Store the retrieved documents per trace with their ranks, and store a reference answer wherever a curated one exists.
- The user question as submitted, before any query rewriting, plus the rewritten query if you use one.
- The retrieved documents in retrieval order, so ordering metrics have something to read.
- The generated response, stored separately from the context it was conditioned on.
- A reference answer where one exists, flagged so the labeled metrics know which traces they can score.
- The retriever and generator configuration version attached to the trace, so a score can be tied back to a knob.
Read the evaluators in triage order
Read retrieval relevance, documents against question, before anything else. A failing verdict there, where the retrieved facts are completely unrelated to the question, means the generator was never given a chance and no prompt change will help.[2] Evidently's left-to-right ordering of retrieval metrics ahead of generation metrics expresses the same priority.[1]
If retrieval relevance passes but groundedness fails, classify the trace as a generation failure: the response contains claims that fall outside the scope of the supplied facts.[2] If both pass and the user is still unhappy, check relevance, response against question, which catches answers that are accurate and grounded yet do not address what was actually asked.[2]
Then work in buckets rather than anecdotes. Sample failing traces, sort them into retrieval misses, grounding failures, and off-question answers, and fix the largest bucket first. Running all the evaluators in the same experiment is what makes this cheap, because one run produces the full attribution instead of one score per investigation.[2]
- Retrieval relevance first: unrelated context makes every downstream score uninterpretable.
- Groundedness next: claims outside the supplied facts are a generator problem, not a retriever problem.
- Response relevance last: grounded and accurate answers can still miss the question.
- One experiment run carrying all evaluators, so attribution comes from a single pass over the sample.
Recall for missing evidence, precision for buried evidence
'Improve retrieval' is not an actionable ticket, so the retrieval metrics have to be read as separate signals about separate knobs. DeepEval's guide maps contextual recall to the embedding model's ability to capture and retrieve relevant information, and contextual precision to the reranker, since precision checks that more relevant nodes are ranked higher than irrelevant ones.[4] Read that way, low recall means the evidence never entered the context window, and low precision with acceptable recall means the evidence is in the window but buried below noise.
Keep both, because dropping recall makes the rest of the panel gameable. Confident AI states the failure directly: without contextual recall, anyone can achieve perfect contextual relevancy by retrieving as little information as possible.[5] A retriever that returns one tightly matched chunk and nothing else scores beautifully on relevancy while silently dropping half the answer.
Vocabulary differs across vendors for the same two axes, which matters when you compare tooling. Patronus AI calls context relevance also known as context precision, and describes context sufficiency as similar to context recall, requiring comparison with a gold answer; it attributes retrieval breakdowns to low-quality embeddings, improper chunking, or insufficient data in the index.[6] When two tools disagree about a metric name, check which artifacts it compares before assuming it measures the same thing.
- Low recall: treat as an embedding-model and coverage signal, including domain vocabulary the current model does not capture.
- Low precision with acceptable recall: treat as a reranker or ranking-logic signal.
- Low recall with an index gap: the document may simply not be there, which no retriever change fixes.
- Metric names mapped per tool to the artifact pair they compare, not to the vendor's label.
Contextual relevancy is the chunk-size and top-K dial
The third retrieval metric answers a different question than the first two. DeepEval maps contextual relevancy to chunk size and top-K, because it measures whether retrieval returns relevant information without much irrelevancy.[4] A context window padded with marginal chunks is the signature failure: the needed facts are present, ranked acceptably, and surrounded by text that dilutes them.
Its other property is operational. Contextual precision and contextual recall need an expected output as ground truth, while contextual relevancy does not.[4] That makes relevancy the retrieval metric you can run continuously without an annotation budget, and precision and recall the ones that depend on a curated set. Run all three rather than picking one, since each covers a different change you might make and a single retrieval score cannot distinguish missing evidence from badly ordered or diluted evidence.
- Tune chunk size and top-K against relevancy, not against the aggregate answer score.
- Keep relevancy as the reference-free retrieval check that runs without labels.
- Pair every relevancy change with a recall read, so shrinking the context does not hide a coverage loss.
Faithfulness is a model setting, response relevance is a prompt setting
Generation-side complaints split into two kinds: the model states things that are not in the sources, or the answer is supported but does not answer the question. The documented metrics separate them and point at different remedies. DeepEval maps faithfulness to the LLM choice and temperature, since the metric checks that the model does not hallucinate and does not contradict factual information in the retrieval context, and maps answer relevancy to the prompt template, checking whether the prompt instructs the model to produce relevant and helpful output from the retrieval context.[4]
Confident AI restates the same division in operational terms: faithfulness measures the hallucination rate of your LLM, while answer relevancy is a direct measure of how well the model follows instructions in the prompt template.[5] Patronus AI labels its hallucination metric as also known as faithfulness, and separates answer relevance, whether the output addresses the user input, from answer correctness against a gold reference.[6] Those are three distinct tickets: a model or decoding change, a prompt-template change, and a correctness gate that needs labels.
One logging detail decides what you can run today. Faithfulness needs the retrieval context, while answer relevancy needs only the input and the output, so relevance checks still work on traces where context logging is incomplete.[4] For format, tone, or language requirements, DeepEval points to a custom criteria metric such as G-Eval, since those are generation-layer concerns the generic metrics do not cover.[4]
- Faithfulness regression: route to model choice, temperature, and grounding instructions.
- Response-relevance regression: route to the prompt template and its instruction wording.
- Correctness regression against a gold answer: route to whichever layer the other metrics already indicted.
- Format, tone, and language rules: cover with a custom criteria judge rather than reading them into faithfulness.
Completeness to context catches the half-answer
There is a generation failure the usual pair misses. Evidently adds answer completeness to context, which asks whether the response makes full use of the relevant information retrieved, and notes it as a common break on multi-source queries where the system returns only part of the answer.[1] Such a trace passes faithfulness, because every claim is supported, and can pass response relevance, because the answer is on topic.
Treat it as a generation-layer check, not a retrieval one: the evidence was retrieved and then under-used. Add it specifically to multi-hop and multi-source question sets, where the gap between what the context contained and what the answer used is widest, and keep it reference-free so it can run on live traffic alongside faithfulness.[1]
- Score completeness on multi-hop and multi-source questions, where partial answers hide behind passing faithfulness.
- Compare what the context supported against what the answer used, rather than against a gold answer.
- Route completeness failures to prompt and synthesis logic, since the evidence was already in the window.
Decide what runs on live traffic and what needs labels
Continuous monitoring forces a clean line between reference-free and labeled metrics. Evidently describes faithfulness and the completeness checks as reference-free, so they run on live traffic without labels to detect hallucinations and degraded performance, and it states the hard constraint on the retrieval side: with post-hoc labeling you cannot compute recall, because you have not defined what should have been retrieved.[1] That is a definitional limit, not a tooling gap, and no amount of production logging removes it.
The vendor splits agree. DeepEval names the reference-free subset explicitly as answer relevancy, faithfulness, and contextual relevancy, and notes that precision and recall need an expected output as ground truth.[4] LangSmith's tutorial shows the same division: only correctness depends on labeled data, while relevance, groundedness, and retrieval relevance run on unlabeled production-style traffic.[2]
So run two tiers. Deploy the reference-free set continuously, including retrieval relevance as a documents-against-question check, and keep a small curated ground-truth set for correctness, context precision, and context recall as a pre-release gate rather than a streaming metric. That split also tells you which regressions you can even see in production: a recall collapse after an index change will not appear in live monitoring, which is a good reason to gate index changes on the labeled set.
- Continuous, reference-free: faithfulness, completeness to context, contextual relevancy, response relevance, retrieval relevance.
- Pre-release, labeled: correctness, context precision, context recall.
- Index and embedding changes gated on the labeled tier, since recall is invisible to production logs.
- The curated set versioned and frozen long enough that configuration comparisons stay meaningful.
Synthesize the eval set when annotation is the blocker
If nobody can annotate gold answers, the labeled tier does not have to stay empty. The Hugging Face cookbook demonstrates the alternative path: generate synthetic factoid question–answer pairs from your own corpus, then filter them with three independent critic agents that score groundedness, relevance, and stand-alone quality on a scale of 1 to 5, keeping only the pairs that score 4 or above on all three.[7] The filtering is the part that makes the set usable, since unfiltered generated questions include items that are unanswerable from the corpus or incomprehensible out of context.
Plan the volume accordingly. Roughly half the generated samples are removed by the critics, so over-generate, and size the surviving set to separate configurations rather than to produce a handful of demo items.[7] A synthetic set of that kind supports an end-to-end correctness score across configuration ablations, which is often the more economical diagnostic when per-layer labels are out of reach: hold the corpus fixed, change one knob at a time, and read which change moves the score.
Keep one judging discipline across both tiers. Require judge prompts to emit the rationale before the verdict or score, the pattern used in both the LangSmith evaluators and the cookbook critics, so a score carries the reason it was assigned and a reviewer can check it.[2][7]
- Generate question–answer pairs from your own corpus, not from a generic benchmark.
- Filter on groundedness, relevance, and stand-alone quality before trusting any score computed from the set.
- Over-generate, because critic filtering removes a large share of the candidates.
- Rationale emitted before the verdict in every judge prompt, on both the live and labeled tiers.