vectory
Back to field notes

Vectory field note

Which RAG evaluation metric names the failing layer

Map each standard RAG metric to the layer it indicts: retrieval relevance and context recall for the retriever, precision for ranking, faithfulness and response relevance for the generator.

October 20269 min read

Key ideas

  • Every diagnostic RAG metric is a comparison between two logged artifacts, so a score attributes to a layer only when you log the question, the retrieved documents, the response, and the reference answer separately.
  • Low context recall is an embedding and coverage signal, low context precision with acceptable recall is a ranking signal, and low contextual relevancy is a chunk-size and top-K signal — one aggregate retrieval score cannot tell them apart.
  • Faithfulness regressions route to model configuration and grounding instructions, while response-relevance regressions route to the prompt template, because that metric reads instruction-following rather than evidence quality.

Workflow

01

Log the four artifacts

02

Attribute the failing layer

03

Pick the retriever knob

04

Split labeled from reference-free

01

Attribute the layer before you change the pipeline

When a RAG assistant returns a wrong answer, the argument that follows is usually about whether to swap the embedding model or rewrite the answer prompt. The current documentation answers that with a two-stage split rather than one aggregate quality score. Evidently's guide frames debugging as explicitly asking whether the system failed to retrieve the right documents or whether the model hallucinated even though the context was correct, and groups its metrics left to right from retrieval to generation.[1] The survey 'Evaluation of Retrieval-Augmented Generation' argues that assessment is hard precisely because of RAG's hybrid structure and its reliance on dynamic knowledge sources, and proposes a unified process that compares quantifiable metrics of the retrieval and generation components separately.[3]

LangSmith's tutorial turns that split into mechanics you can implement. It defines four evaluators, each comparing exactly two artifacts drawn from the question, the response, the retrieved documents, and a reference answer: correctness, relevance, groundedness, and retrieval relevance.[2] The artifact-pair framing is what makes a score attributable to a layer. Retrieval relevance never looks at the answer, so it cannot be contaminated by generation quality, and groundedness never looks at the question, so it cannot be dragged down by a retrieval miss.[2]

The practical consequence is a logging requirement, not a dashboard requirement. If your traces store only the question and the final answer, no metric you add later can separate the layers, because the comparison it needs is missing. Store the retrieved documents per trace with their ranks, and store a reference answer wherever a curated one exists.

  • The user question as submitted, before any query rewriting, plus the rewritten query if you use one.
  • The retrieved documents in retrieval order, so ordering metrics have something to read.
  • The generated response, stored separately from the context it was conditioned on.
  • A reference answer where one exists, flagged so the labeled metrics know which traces they can score.
  • The retriever and generator configuration version attached to the trace, so a score can be tied back to a knob.
02

Read the evaluators in triage order

Read retrieval relevance, documents against question, before anything else. A failing verdict there, where the retrieved facts are completely unrelated to the question, means the generator was never given a chance and no prompt change will help.[2] Evidently's left-to-right ordering of retrieval metrics ahead of generation metrics expresses the same priority.[1]

If retrieval relevance passes but groundedness fails, classify the trace as a generation failure: the response contains claims that fall outside the scope of the supplied facts.[2] If both pass and the user is still unhappy, check relevance, response against question, which catches answers that are accurate and grounded yet do not address what was actually asked.[2]

Then work in buckets rather than anecdotes. Sample failing traces, sort them into retrieval misses, grounding failures, and off-question answers, and fix the largest bucket first. Running all the evaluators in the same experiment is what makes this cheap, because one run produces the full attribution instead of one score per investigation.[2]

  • Retrieval relevance first: unrelated context makes every downstream score uninterpretable.
  • Groundedness next: claims outside the supplied facts are a generator problem, not a retriever problem.
  • Response relevance last: grounded and accurate answers can still miss the question.
  • One experiment run carrying all evaluators, so attribution comes from a single pass over the sample.
03

Recall for missing evidence, precision for buried evidence

'Improve retrieval' is not an actionable ticket, so the retrieval metrics have to be read as separate signals about separate knobs. DeepEval's guide maps contextual recall to the embedding model's ability to capture and retrieve relevant information, and contextual precision to the reranker, since precision checks that more relevant nodes are ranked higher than irrelevant ones.[4] Read that way, low recall means the evidence never entered the context window, and low precision with acceptable recall means the evidence is in the window but buried below noise.

Keep both, because dropping recall makes the rest of the panel gameable. Confident AI states the failure directly: without contextual recall, anyone can achieve perfect contextual relevancy by retrieving as little information as possible.[5] A retriever that returns one tightly matched chunk and nothing else scores beautifully on relevancy while silently dropping half the answer.

Vocabulary differs across vendors for the same two axes, which matters when you compare tooling. Patronus AI calls context relevance also known as context precision, and describes context sufficiency as similar to context recall, requiring comparison with a gold answer; it attributes retrieval breakdowns to low-quality embeddings, improper chunking, or insufficient data in the index.[6] When two tools disagree about a metric name, check which artifacts it compares before assuming it measures the same thing.

  • Low recall: treat as an embedding-model and coverage signal, including domain vocabulary the current model does not capture.
  • Low precision with acceptable recall: treat as a reranker or ranking-logic signal.
  • Low recall with an index gap: the document may simply not be there, which no retriever change fixes.
  • Metric names mapped per tool to the artifact pair they compare, not to the vendor's label.
04

Contextual relevancy is the chunk-size and top-K dial

The third retrieval metric answers a different question than the first two. DeepEval maps contextual relevancy to chunk size and top-K, because it measures whether retrieval returns relevant information without much irrelevancy.[4] A context window padded with marginal chunks is the signature failure: the needed facts are present, ranked acceptably, and surrounded by text that dilutes them.

Its other property is operational. Contextual precision and contextual recall need an expected output as ground truth, while contextual relevancy does not.[4] That makes relevancy the retrieval metric you can run continuously without an annotation budget, and precision and recall the ones that depend on a curated set. Run all three rather than picking one, since each covers a different change you might make and a single retrieval score cannot distinguish missing evidence from badly ordered or diluted evidence.

  • Tune chunk size and top-K against relevancy, not against the aggregate answer score.
  • Keep relevancy as the reference-free retrieval check that runs without labels.
  • Pair every relevancy change with a recall read, so shrinking the context does not hide a coverage loss.
05

Faithfulness is a model setting, response relevance is a prompt setting

Generation-side complaints split into two kinds: the model states things that are not in the sources, or the answer is supported but does not answer the question. The documented metrics separate them and point at different remedies. DeepEval maps faithfulness to the LLM choice and temperature, since the metric checks that the model does not hallucinate and does not contradict factual information in the retrieval context, and maps answer relevancy to the prompt template, checking whether the prompt instructs the model to produce relevant and helpful output from the retrieval context.[4]

Confident AI restates the same division in operational terms: faithfulness measures the hallucination rate of your LLM, while answer relevancy is a direct measure of how well the model follows instructions in the prompt template.[5] Patronus AI labels its hallucination metric as also known as faithfulness, and separates answer relevance, whether the output addresses the user input, from answer correctness against a gold reference.[6] Those are three distinct tickets: a model or decoding change, a prompt-template change, and a correctness gate that needs labels.

One logging detail decides what you can run today. Faithfulness needs the retrieval context, while answer relevancy needs only the input and the output, so relevance checks still work on traces where context logging is incomplete.[4] For format, tone, or language requirements, DeepEval points to a custom criteria metric such as G-Eval, since those are generation-layer concerns the generic metrics do not cover.[4]

  • Faithfulness regression: route to model choice, temperature, and grounding instructions.
  • Response-relevance regression: route to the prompt template and its instruction wording.
  • Correctness regression against a gold answer: route to whichever layer the other metrics already indicted.
  • Format, tone, and language rules: cover with a custom criteria judge rather than reading them into faithfulness.
06

Completeness to context catches the half-answer

There is a generation failure the usual pair misses. Evidently adds answer completeness to context, which asks whether the response makes full use of the relevant information retrieved, and notes it as a common break on multi-source queries where the system returns only part of the answer.[1] Such a trace passes faithfulness, because every claim is supported, and can pass response relevance, because the answer is on topic.

Treat it as a generation-layer check, not a retrieval one: the evidence was retrieved and then under-used. Add it specifically to multi-hop and multi-source question sets, where the gap between what the context contained and what the answer used is widest, and keep it reference-free so it can run on live traffic alongside faithfulness.[1]

  • Score completeness on multi-hop and multi-source questions, where partial answers hide behind passing faithfulness.
  • Compare what the context supported against what the answer used, rather than against a gold answer.
  • Route completeness failures to prompt and synthesis logic, since the evidence was already in the window.
07

Decide what runs on live traffic and what needs labels

Continuous monitoring forces a clean line between reference-free and labeled metrics. Evidently describes faithfulness and the completeness checks as reference-free, so they run on live traffic without labels to detect hallucinations and degraded performance, and it states the hard constraint on the retrieval side: with post-hoc labeling you cannot compute recall, because you have not defined what should have been retrieved.[1] That is a definitional limit, not a tooling gap, and no amount of production logging removes it.

The vendor splits agree. DeepEval names the reference-free subset explicitly as answer relevancy, faithfulness, and contextual relevancy, and notes that precision and recall need an expected output as ground truth.[4] LangSmith's tutorial shows the same division: only correctness depends on labeled data, while relevance, groundedness, and retrieval relevance run on unlabeled production-style traffic.[2]

So run two tiers. Deploy the reference-free set continuously, including retrieval relevance as a documents-against-question check, and keep a small curated ground-truth set for correctness, context precision, and context recall as a pre-release gate rather than a streaming metric. That split also tells you which regressions you can even see in production: a recall collapse after an index change will not appear in live monitoring, which is a good reason to gate index changes on the labeled set.

  • Continuous, reference-free: faithfulness, completeness to context, contextual relevancy, response relevance, retrieval relevance.
  • Pre-release, labeled: correctness, context precision, context recall.
  • Index and embedding changes gated on the labeled tier, since recall is invisible to production logs.
  • The curated set versioned and frozen long enough that configuration comparisons stay meaningful.
08

Synthesize the eval set when annotation is the blocker

If nobody can annotate gold answers, the labeled tier does not have to stay empty. The Hugging Face cookbook demonstrates the alternative path: generate synthetic factoid question–answer pairs from your own corpus, then filter them with three independent critic agents that score groundedness, relevance, and stand-alone quality on a scale of 1 to 5, keeping only the pairs that score 4 or above on all three.[7] The filtering is the part that makes the set usable, since unfiltered generated questions include items that are unanswerable from the corpus or incomprehensible out of context.

Plan the volume accordingly. Roughly half the generated samples are removed by the critics, so over-generate, and size the surviving set to separate configurations rather than to produce a handful of demo items.[7] A synthetic set of that kind supports an end-to-end correctness score across configuration ablations, which is often the more economical diagnostic when per-layer labels are out of reach: hold the corpus fixed, change one knob at a time, and read which change moves the score.

Keep one judging discipline across both tiers. Require judge prompts to emit the rationale before the verdict or score, the pattern used in both the LangSmith evaluators and the cookbook critics, so a score carries the reason it was assigned and a reviewer can check it.[2][7]

  • Generate question–answer pairs from your own corpus, not from a generic benchmark.
  • Filter on groundedness, relevance, and stand-alone quality before trusting any score computed from the set.
  • Over-generate, because critic filtering removes a large share of the candidates.
  • Rationale emitted before the verdict in every judge prompt, on both the live and labeled tiers.

Use this checklist

Before you ship

Sources

Watch and read the original material

  1. Evidently AI, LLM evaluation guide: RAG evaluation

    Source for framing debugging as retrieval failure versus hallucination despite correct context, grouping metrics from retrieval to generation, answer completeness to context as a separate generation failure on multi-source queries, faithfulness and completeness as reference-free metrics for live traffic, and the constraint that post-hoc labeling cannot produce recall. Read 7 October 2026.

  2. LangSmith docs, Evaluate a RAG application tutorial

    Source for the four evaluators, correctness, relevance, groundedness, and retrieval relevance, each comparing exactly two artifacts drawn from question, response, retrieved documents, and reference answer; for only correctness depending on labeled data; and for judge prompts emitting the rationale before the verdict. Read 7 October 2026.

  3. Evaluation of Retrieval-Augmented Generation: A Survey (arXiv:2405.07437)

    Source for the argument that RAG assessment is hard because of its hybrid structure and reliance on dynamic knowledge sources, and for the proposed unified process that compares quantifiable metrics of the retrieval and generation components separately. Read 7 October 2026.

  4. DeepEval, RAG evaluation guide

    Source for mapping contextual precision to the reranker, contextual recall to the embedding model, contextual relevancy to chunk size and top-K, faithfulness to LLM choice and temperature, and answer relevancy to the prompt template; for the reference-free subset of answer relevancy, faithfulness, and contextual relevancy; for precision and recall requiring an expected output; and for custom criteria metrics such as G-Eval. Read 7 October 2026.

  5. Confident AI, RAG evaluation metrics: answer relevancy, faithfulness and more

    Source for the warning that without contextual recall anyone can achieve perfect contextual relevancy by retrieving as little information as possible, and for faithfulness measuring hallucination rate while answer relevancy measures how well the model follows instructions in the prompt template. Read 7 October 2026.

  6. Patronus AI, RAG evaluation metrics

    Source for context relevance also being called context precision, context sufficiency being similar to context recall and requiring comparison with a gold answer, retrieval breakdowns attributed to low-quality embeddings, improper chunking, or insufficient data, and the hallucination metric also being known as faithfulness alongside separate answer relevance and answer correctness. Read 7 October 2026.

  7. Hugging Face Open-Source AI Cookbook, RAG evaluation

    Source for generating synthetic factoid question-answer pairs from your own corpus, filtering them with three independent critic agents scoring groundedness, relevance, and stand-alone quality on a 1 to 5 scale while keeping only pairs scoring 4 or above on all three, roughly half of generated samples being filtered out, and critics emitting a rationale before the score. Read 7 October 2026.

Evaluation guides

Go deeper on what to evaluate

Vectory uses essential cookies and optional analytics to improve the site. You can update choices any time in cookie preferences. Privacy Policy

Customize