Shipping an LLM feature is easy. Knowing whether it still works six months, three model versions, and two prompt refactors later is the hard part. Over the last few years of building LLM-backed systems that serve real traffic, I've converged on a framework that actually catches regressions before users do — without requiring a research lab's budget.
Start with a golden dataset, not vibes
Every eval practice starts with the same artifact: a set of input-output pairs that encode "this is what good looks like." The mistake I see most often is treating the golden dataset as a one-time artifact assembled by an engineer in an afternoon. A useful golden set is grown, not written.
The highest-signal examples come from production incidents and user corrections. Every time someone flags a bad response, fixes an answer, or rephrases a query after getting a bad result, that interaction is a candidate for the dataset — with the corrected output as the expected answer. This keeps your evals anchored to the failures your users actually experience rather than the failures you imagine.
Aim for breadth over depth early: 100–200 examples covering your distinct task types (extraction, summarization, classification, grounded QA, open-ended generation) beats 1,000 near-duplicate QA pairs. Tag each example with its task type and difficulty, because aggregate pass rates lie — a model can hold 92% overall while collapsing on one task type, and you won't notice unless you slice.
One more thing: version your golden set. When you add an example because a prompt change fixed a real incident, tag it with the date and the incident. Six months later, when someone asks "why does this prompt look weird," the answer is in the dataset history.
LLM-as-judge works — if you know what it's lying about
Human grading doesn't scale, so most teams graduate to LLM-as-judge: a stronger model scores outputs against a rubric. This works surprisingly well for structured tasks — our judge-based scores correlated with human ratings at roughly 0.8 on extraction and summarization tasks. But the judge has systematic biases, and if you don't correct for them, your evals will quietly certify bad behavior.
Self-preference bias. A judge model favors outputs that resemble its own generation style. If your judge is from the same model family as your generator, expect inflated scores. Mitigate it: use a different model family for judging than for generation, and periodically calibrate against human labels on a held-out slice. If judge-human agreement drifts below your threshold (we use 0.75 correlation as the floor), the judge needs recalibration, not the pipeline.
Verbosity bias. Judges reward longer answers, even when the rubric says "be concise." This is one of the most replicated findings in eval research, and it bites hardest on summarization evals — the judge will prefer the 8-sentence summary over the correct 3-sentence one. Counter it with explicit rubric instructions ("penalize answers that include information not requested"), and add a length-normalized check: flag any output more than 2x the reference length for human review.
Position bias. When the judge compares two outputs (A vs. B), it favors whichever came first — or second, depending on the model. The fix is trivial and widely skipped: run every pairwise comparison in both orders and only count wins that are consistent across both. Inconsistent orderings are ties. Yes, this doubles judge cost. It's still cheaper than shipping a "better" prompt that wasn't actually better.
Also worth internalizing: judges are worst at exactly the tasks where you most want them — open-ended generation with subjective quality bars. Use LLM judges for structured, rubric-checkable tasks, and reserve humans for the fuzzy ones.
Human-in-the-loop: sample like you mean it
You can't have humans review everything, so sampling strategy is the whole game. Random sampling wastes human attention on easy cases; the defects live in the tails.
Uncertainty sampling works well in practice: have the judge emit a confidence score alongside each grade, and route the lowest-confidence grades to humans. In our experience, the bottom 10% of judge-confidence outputs contained the majority of real defects — a 10x leverage on reviewer time.
Stratified sampling guards against blind spots: sample a fixed minimum per task type, per language, per user cohort — even when volumes are low. The bug that only affects 0.3% of traffic can still be the one that ends up in a support ticket from your largest customer.
And always keep a canary slice: 20–50 production queries per day that humans review no matter what. This is your ground truth against which you validate the judge itself. If judge-human agreement on the canary slice drops, you have a judge problem, and every eval result since the last healthy check is suspect.
Regression-test your prompts like code
A prompt change is a code change. Treat it like one: every pull request that touches a prompt, a model version, or a retrieval config should run the eval suite and show the diff in scores — per task type, not just aggregate.
This catches a class of failure I call capability whack-a-mole: you fix the model refusing to answer financial questions, and it starts refusing to answer legal questions. Aggregate accuracy stays flat; per-slice accuracy reveals the trade. Without per-slice regression gates, prompt engineering degrades into moving bugs around.
Pin your model versions in the eval harness. Model providers update weights under the same version string more often than they'd like you to know. Record the actual model identifier (and, when available, the system fingerprint) in every eval run so a mysterious score shift is attributable. When a pinned model degrades without any change on your side, that's a provider-side change — and it's worth knowing before your users tell you.
Set thresholds as gates, not decorations. A suite that "runs" but never blocks a merge is theater. We gate on two conditions: no task-type slice may drop more than 2 absolute points, and the canary agreement floor must hold. Everything else is informational.
Track evals in CI, and make the history visible
Evals that only run locally get forgotten. Wire them into CI so every prompt or model change produces a score diff, and store the results in a time-series store — not a spreadsheet, not a Slack thread. You want to be able to answer "when did summarization quality start declining?" with a chart, not an archaeology expedition.
The metrics worth trending: per-task-type pass rates, judge-human agreement on the canary slice, output length distributions (a sudden shift often precedes a quality shift), and cost per eval run (judge calls are real money; a suite that costs $40 per run gets run less often than one that costs $4).
A few operational numbers from experience: a solid starting suite is 150–300 golden examples with judge grading at roughly $0.01–0.05 per example per run — cheap enough to run on every prompt PR. Human review budget goes entirely to the canary slice plus uncertainty-sampled outliers, typically a few dozen items per week for a mid-size feature. That ratio — broad automated coverage, narrow deep human coverage — is the whole framework in one sentence.
The bottom line
Eval quality compounds. The teams that do this well aren't the ones with the fanciest judge prompts; they're the ones whose golden datasets grow from production incidents, whose judges are calibrated against humans on an ongoing basis, and whose prompt changes face the same regression discipline as code changes. Start with a golden set built from real failures, add a calibrated judge with known biases corrected, sample humans where the judge is weakest, and gate your merges on per-slice scores. Everything else is optimization.