LLM-as-Judge Is Not a Score. It Is a Reasoning Contract.

LLM-as-Judge Is Not a Score. It Is a Reasoning Contract.

Most teams use an LLM judge as a scorer: prompt it, read the number, gate on it. That is the mistake. A scorer returns a number you cannot audit and cannot trust. A judge is a contract with four clauses: the evidence it is allowed to use, the rubric that binds its reasoning, the schema it must answer in, and the external check that can veto it. Get those four right and the number means something. Skip them and you have automated a vibe.

October 6, 2026
Harrison Guo
7 min read
AI Engineering Evaluation

Most teams use an LLM judge the same way. Write a prompt that ends in “rate this from 1 to 10”, call the model, read the number, gate on it. Ship.

That is not a judge. That is a vibe with a number stapled to it.

A judge, the kind you can actually build automation on, is not a scoring function. It is a contract. It specifies what the judge is allowed to know, what standard binds its reasoning, what shape its answer must take, and what can overrule it. Get those four clauses right and the number at the end means something. Skip them, which is the default, and you have automated the feeling of rigor without any of it.

Here are the four clauses, and what goes wrong when each is missing.

The bare number has nothing holding it up

Start with what you get from the naive approach, because it fails in three ways at once and they are worth separating.

You cannot audit it. A 7 arrives with no reason attached. When a stakeholder asks why this output scored a 7 and that one a 4, you have nothing. The judge did reasoning, then threw it away and handed you the digit.

You cannot trust it over time. The same output scored by the same prompt drifts when the model updates, and you will not notice, because nothing pins today’s 7 to last week’s 7. I wrote a whole piece on measuring exactly this kind of drift in a model’s own confidence numbers; a judge’s scores drift the same way and for the same reasons.

You cannot tell if it is even right. The judge might have scored the output a 9 because it believed the output contained something it does not contain. The number looks fine. The reasoning under it was false. Nothing caught that.

Every one of these is fixed by a clause. None of them is fixed by a better prompt.

Clause one: the evidence it is allowed to use

A judge reasons over inputs. The first clause of the contract says exactly what those inputs are, and just as importantly, what they are not.

The rule that saves you the most pain: the judge must never treat the producer’s own claims as evidence of the output. If the thing you are evaluating came with a manifest, a plan, a self-description, the judge cannot read “the changelog says tests were added” as proof that tests exist. It has to look at the output. This sounds obvious and it is violated constantly, because the producer’s claims are right there in the payload and they are easy to read. A judge that grades the claim instead of the artifact will pass a pipeline that executes a wrong plan perfectly.

The second half of this clause: where a fact can be computed, compute it, and feed it in as verified evidence. Do not ask the judge to count things by eye that a function can count exactly. The judge should receive “section count: 5, every link resolves, no banned term present” as established fact and reason on top of it. It reasons over the facts. It does not re-derive them by vibe.

Clause two: the rubric binds the reasoning

A number needs a standard, and for anything subjective the standard cannot be a bare absolute scale. “Rate the quality from 1 to 10” gives the judge nothing to anchor to, so it anchors to nothing, and 7 means whatever it felt like this call.

Replace the absolute with a comparison against anchors. Not “how good is this, 1 to 10” but “here is a known-good example and a known-borderline example; where does this sit relative to them”. Anchored comparison is reproducible in a way an absolute score never is, because the judge is measuring a distance from a fixed point instead of inventing a point each time.

And stop reporting fine-grained scores you cannot defend. If your judge cannot reliably tell a 6 from a 7, do not emit a 1-to-10 scale. Emit the bands you can actually defend: fails, borderline, clears. A coarse score you can stand behind beats a precise one you made up. This is the same argument as a wrong ruler being worse than no ruler: a precise fake number does more damage than an honest coarse one, because people trust the decimal places.

Clause three: a schema that must name the failure

The judge’s output is part of the contract, and “a number” is not an output schema.

Require structured output that names what is wrong, not just how wrong. A verdict is a label, the evidence it rests on, and the specific failure when it does not pass. “Reject, because the request asked for a refund and the reply refuses one, see transcript span 3” is a verdict you can act on, audit, and aggregate. “4” is not.

This also makes the next clause possible. A judge that has to state the evidence its score rests on is a judge whose evidence you can check. A judge that only emits a number has hidden the one thing you need to verify it.

Clause four: the rule is the lie detector

This is the clause that turns the first three into something trustworthy, and it is the one almost nobody has.

A subjective score may not contradict an objective fact. If the judge says the answer is excellent and cites “every required field is present” as a load-bearing reason, and your deterministic check says a required field is missing, the score is vetoed. Not averaged. Not softened. Vetoed. The judge’s claimed evidence was false, so its verdict is void.

This is an authority split. The model is allowed to reason about quality. It is not allowed to be the final word on facts. Facts are the rule engine’s job, and the rule engine gets to overrule the model whenever the model’s reasoning depends on a fact that is false. The judge proposes; the rule disposes. Without this clause, a confident hallucination becomes a confident score and flows straight into your gate.

This is also where the technique boundary lives in practice. Deciding what the rule computes versus what the judge reasons about is not a style choice. It is the line that decides whether a false fact can move a score.

The contract drifts, so you test it like code

A contract you sign once and never check is a contract the other party stops honoring. The judge drifts: the model updates, someone tweaks the prompt, and the scores move while you are looking elsewhere.

So you pin it. Keep a fixed, versioned set of cases you have already labeled by hand, and re-run the judge against them like a regression suite. When a change moves the scores on that frozen corpus, you see it before production does. Report the disagreements between the judge and the labels instead of hiding them in an average, because the disagreements are where the judge is about to fail. A judge without a regression corpus is a test you wrote and never ran.

The one sentence to take

An LLM judge is not a function that returns a score. It is a contract: these are the facts you may use, this is the standard you are held to, this is the shape of the answer you must give, and this is the check that overrules you when your reasoning rests on something false.

Write that contract and the number at the end is worth gating on. Skip it, prompt the model for a 1-to-10, and read the digit, and you have not measured anything. You have measured the harness and called it the model, one abstraction up. The score was never the judge. The contract is.

🎧 More Ways to Consume This Content

I occasionally advise small teams on backend reliability, Go performance, and production AI systems. Learn more: /services

Comments

This space is waiting for your voice.

Comments will be supported shortly. Stay connected for updates!

Preview of future curated comments

This section will display user comments from various platforms like X, Reddit, YouTube, and more. Comments will be curated for quality and relevance.