vectory
Back to field notes

Vectory field note

Calibrate an LLM judge before you trust its score

A held-out calibration procedure for an LLM judge: de-noised human labels, a published agreement number, order-swap and length-inflation probes, and an abstention label that cannot read as a pass.

September 20269 min read

Key ideas

  • A judge score is evidence only when you can name the metric, the held-out set it was measured on, and the human-human agreement ceiling it sits under.
  • Position and verbosity bias are measurable on your own data: swap the order of every pairwise comparison, and re-score padded and compressed copies of responses you already labeled.
  • Give the judge an explicit abstention label, report its rate as its own metric, and keep abstentions out of the pass rate instead of letting them read as passes.

Workflow

01

Label a held-out set

02

Report the agreement number

03

Probe order and length bias

04

Fix the scale and abstention

01

Treat the judge as a supervised model, not a prompt

The number your judge reports is a model output, and it becomes evidence only when you can name the labeled set it was measured against. Evidently's guide puts the order plainly: label the dataset manually first, and reserve part of it as a held-out set used only to test the final prompt.[2] The Hugging Face cookbook shows how small that set can be to start, sampling 7 items per score level for 28 balanced examples and noting that roughly 30 examples is enough for a first read on how a judge is performing.[1] Weight the sample toward the failure modes you care about rather than toward the traffic mix you see most often.

Then deal with label noise before you blame the judge for it. On the feedbackQA data used in the cookbook, the two human annotators correlate at only 0.563 Pearson, which the cookbook reads as noise in the ground truth itself: you cannot expect any algorithmic evaluation to come that close to it.[1] Its fix is to keep only the rows where both raters agree and then rebalance across score levels, so the calibration set is clean and no single label dominates.[1]

Split what is left in two. An iteration half is where you read disagreements, rewrite the judge prompt, and read again. A held-out half is touched once, at the end, and produces the number you publish.[2] Tune against the held-out half and you no longer have a calibration result, you have a fit.

  • Written definitions for every label, plus a one-line rationale on each labeled item.
  • A second annotator's labels on the same items, stored alongside the first set.
  • A de-noised set rebalanced across score levels, so no label dominates the result.
  • An iteration split for prompt work and a held-out split touched once, at the end.
  • The judge model and prompt version recorded with every calibration run.
02

Publish the agreement number and the ceiling above it

Pick the metric before the run, because the metric decides what counts as good. Evidently recommends precision and recall for binary verdicts, and points out that recall is the one to optimize when the judge exists to catch bad outputs.[2] Langfuse's method for the same question is to compare judge scores against human annotation scores on the same data and quantify the gap with Cohen's Kappa and a confusion matrix.[4] For a graded scale, correlation against human labels is the readable number: the cookbook reports a naive judge prompt at 0.567 correlation and a redesigned prompt at 0.843 on the same items.[1]

Set the acceptance bar before you look at the result. Arize publishes an explicit one: 75-90% label agreement with humans is the point at which you are ready to scale the judge.[3] Write your own threshold down first, so a near miss becomes a prompt revision rather than a rounding decision made under release pressure.

Then report the ceiling next to the score. Because the labels carry annotator disagreement, at 0.563 Pearson between the two feedbackQA raters, no judge can reach perfect agreement against them, and a reviewer expecting 100% is reading the wrong target.[1] Publishing the human-human agreement beside the judge-human agreement shows how much of the gap is actually available to close.

  • One named agreement metric per verdict type, chosen before the held-out run.
  • Recall weighted over precision when the judge exists to catch bad outputs.
  • A confusion matrix stored with the run, not only the headline agreement figure.
  • Human-human agreement published beside judge-human agreement.
  • The acceptance bar written down before anyone reads the held-out result.
03

Run every pairwise comparison in both orders

A pairwise win rate is only as stable as the order you fed the candidates in. The LLM-as-a-Judge survey reports a 2023 study in which GPT-4, the strongest model tested, gave the same verdict both ways in only about two-thirds of cases, and a later large-scale analysis characterized the effect as systematic rather than a matter of chance, persisting across models.[5] Confident AI describes the same tendency from the practitioner side: judges generally prefer the first generated output over the second one.[6]

The correction is cheap and mechanical. Run every comparison twice with the two responses swapped, record both verdicts instead of only the aggregate, and count a win only when the same response prevails in both orders.[5] Arize gives the same instruction as an evaluation-design rule: randomize candidate order and evaluate both permutations.[3] Report the share of pairs that agree across orders as a first-class calibration statistic, and send the pairs that flip to a tie bucket or to human review rather than letting the ordering break the tie for you.

If the consistency rate is too low to act on, two documented routes remain. Order swapping combined with majority voting across repeated runs was found to work in one survey where common improvement strategies are not fully effective, and few-shot in-context examples lifted GPT-4 consistency from 65.0% to 77.5% at the cost of more input tokens.[5][6] The other route is to stop comparing: Evidently suggests scoring each response directly against named criteria, which removes the ordering channel entirely.[2] That changes the metric definition, so re-validate against your held-out labels afterwards.

  • Both orderings run and stored for every pairwise comparison.
  • An order-consistency rate reported beside the win rate.
  • Wins counted only when the same response prevails in both orders.
  • Flipped pairs routed to a tie bucket or to human review.
  • Consistency re-measured after any judge model or prompt change.
04

Probe verbosity with a length-inflation test

Verbosity bias has an experimental signature you can reproduce on your own data. When responses were rewritten at greater length without adding new information, Claude and GPT-3.5 chose the longer one more than 90% of the time.[5] Confident AI reports the same pattern and its consequence: a judge that prefers verbose text over concise text produces scores that stop tracking output quality.[6]

The starkest version of that failure is a fixed null response, ignoring the input entirely, still reaching high win rates on automatic benchmarks, which the survey presents as evidence that such scores do not track quality.[5] If your model got wordier this quarter and your judge score rose, you have not yet separated those two explanations.

So build the probe. Take held-out responses the judge already scored, pad each one with restatement and hedging that adds no new information, re-score, and report the share where the padded version scores higher. Then run the mirror test: compress responses to remove padding while preserving content, and check whether scores drop. A judge that moves in both directions is tracking length, not substance. Correlating judge score against response token count across the held-out set gives you the same estimate as a single figure.

Fix it at the rubric level first. Anchor each score level to substantive properties, the way the cookbook defines its top level as relevant, direct, detailed, and addressing all the concerns raised in the question, and pass a reference answer when one exists.[1] Confident AI reports that reference-based scoring, where an expected output anchors the judgment, improves calibration and reduces variability on criteria such as factual correctness.[6] For leaderboard-style comparisons, length-controlled win rates of the kind used in AlpacaEval 2.0 are the documented control, and publishing mean response length beside every judge score keeps a length shift visible.[5]

  • A padded rerun of held-out responses, with the share of length-driven increases reported.
  • A compression rerun as the mirror test, checking that scores fall when padding is removed.
  • Judge score correlated against response token count on the held-out set.
  • A reference answer passed to the judge wherever one exists.
  • Mean response length published beside every judge score.
05

Collapse the scale and anchor every level

Scale design turns out to be one of the largest single levers on agreement. The cookbook replaced a 0-10 float prompt with a small integer scale, either 1-4 or 1-5, an anchored description for each level, and an Evaluation field emitted before the rating; correlation with human scores moved from 0.567 to 0.843.[1] That was a prompt change, not a model change.

Evidently argues for binary or coarse labels on the same grounds: two-way choices are more reliable and consistent than deciding whether politeness scores 73 or 82, because LLMs are not naturally calibrated for high-precision scoring.[2] It offers a test you can apply to your own rubric today, which is that if you cannot articulate the difference between a 3 and a 4, the scale should be collapsed.[2] Arize arrives at the same place from stability: binary outputs tend to produce more stable and reliable evaluations, graded numeric scores are reserved for comparing prompt versions, and the explanation should be generated before the label.[3]

Where a holistic judgment still feels arbitrary, decompose it. Confident AI describes breaking an output into close-ended yes/no judgments so the score is computed from a count rather than produced as an impression.[6] The same discipline applies to multi-dimensional criteria: split them into separate single-criterion judges and combine the results by explicit rules, rather than asking one prompt to weigh helpfulness, grounding, and format at once.

  • A binary or three-way verdict by default, with a written definition for every label.
  • A small anchored scale, 1-4 or 1-5, if you keep a graded score at all.
  • Reasoning emitted before the verdict, with both fields required to be populated.
  • One criterion per judge, combined downstream by explicit rules.
  • Low temperature, with repeated runs aggregated by majority vote.
06

Give the judge a way to say it cannot tell

A judge with no way to abstain will guess, and the guess lands in your pass rate. Evidently recommends an explicit middle or opt-out category such as partially relevant or unknown, which avoids forcing the model to make a decision without sufficient data.[2] Define exactly when the judge should use it, or borderline cases keep collapsing into the nearest verdict.

Make it a real category in the platform rather than a string you parse after the fact. Langfuse distinguishes numeric, categorical, and boolean score types and lets you declare the allowed categories for a categorical score, so abstention becomes a value the system knows about.[4] Arize's rule that the explanation is generated before the label supplies the other half of the audit trail, because every abstention then carries the reason the judge could not decide.[3]

Decide the downstream policy before you ship. Report the abstention rate as its own metric, exclude abstentions from the pass rate rather than counting them as passes, and route them to human review. An abstention rate that climbs after a retrieval change is one of the more useful signals a judge can give you, and folding it into a single pass number hides exactly the cases worth reading.

  • An explicit abstention label with a written trigger condition.
  • A declared categorical score type, so abstention is not a parsed string.
  • Abstention rate reported as its own metric on every run.
  • Abstentions excluded from the pass rate and routed to human review.
07

Recalibrate whenever the judge changes

Calibration belongs to a judge version, not to a judge. Prompts do not transfer cleanly between models, so a model swap, a prompt edit, or a rubric tweak invalidates the agreement number you published and needs the held-out run again. Langfuse versions evaluator definitions and has active evaluation rules pick up the latest version automatically, which makes the version boundary explicit enough to hang a recalibration trigger on.[4]

Keep the limits of the result in view as well. Agreement is measured against labels that already contain annotator disagreement, so the achievable ceiling sits below perfect agreement and a target of 100% is not a target at all.[1] Aggregate agreement can also look healthy while individual items disagree heavily, which is why the confusion matrix and the list of order-flipped pairs are more useful for debugging than the headline figure.

Finally, treat the calibration set as a maintained artifact. When a new failure mode shows up in production, label it and add it to the iteration split, and leave the held-out split frozen long enough that version-to-version comparisons still mean something. A judge whose calibration set never changes will keep agreeing with humans about last quarter's failures.

  • A recalibration trigger on every judge model, prompt, and rubric version change.
  • The evaluator version recorded with each published agreement number.
  • New production failure modes labeled into the iteration split, not the held-out split.
  • The held-out split frozen long enough for version-to-version comparison.

Use this checklist

Before you ship

Sources

Watch and read the original material

  1. Hugging Face Cookbook, Using LLM-as-a-judge for an automated and versatile evaluation

    Source for the 0.563 Pearson correlation between the two feedbackQA human annotators, de-noising by keeping only rows where both raters agree, 7 items per score level for 28 balanced examples, roughly 30 examples for a first read, the move from a 0-10 float prompt to a small 1-4 or 1-5 integer scale with an Evaluation field before the rating, the anchored top-level rubric wording, and the correlation change from 0.567 to 0.843. Read 28 September 2026.

  2. Evidently AI, LLM-as-a-judge: a complete guide

    Source for labeling the dataset manually first and holding out a split used only to test the final prompt, precision and recall for binary verdicts with recall favored when catching bad outputs, binary and coarse labels being more reliable than high-precision numeric scoring, the 3-versus-4 collapse test, direct per-response scoring as an alternative to pairwise, and the partially relevant or unknown opt-out category. Read 28 September 2026.

  3. Arize AI, LLM as a Judge

    Source for the 75-90% label-agreement bar for scaling a judge, randomizing candidate order and evaluating both permutations, binary outputs producing more stable and reliable evaluations, graded numeric scores reserved for prompt-version comparisons, and generating the explanation before the label. Read 28 September 2026.

  4. Langfuse Docs, LLM-as-a-Judge evaluation methods

    Source for comparing judge scores against human annotation scores on the same data with Cohen's Kappa and a confusion matrix, the numeric, categorical, and boolean score types with declared allowed categories, and versioned evaluator definitions that active rules pick up automatically. Read 28 September 2026.

  5. Wikipedia, LLM-as-a-Judge

    Source for the 2023 finding that GPT-4 gave the same pairwise verdict both ways in about two-thirds of cases and the later analysis calling position bias systematic rather than chance, Claude and GPT-3.5 preferring length-inflated responses more than 90% of the time, the null response reaching high win rates on automatic benchmarks, order swapping with majority voting, and length-controlled win rates as used in AlpacaEval 2.0. Read 28 September 2026.

  6. Confident AI, Why LLM-as-a-Judge is the best LLM evaluation method

    Source for judges preferring the first generated output over the second and verbose text over concise text, few-shot in-context examples lifting GPT-4 consistency from 65.0% to 77.5% at the cost of more input tokens, reference-based scoring improving calibration on criteria such as factual correctness, and decomposing an output into close-ended yes/no judgments for a formula-backed score. Read 28 September 2026.

Vectory uses essential cookies and optional analytics to improve the site. You can update choices any time in cookie preferences. Privacy Policy

Customize