SYNTRX.ai
Blog

LLM Hallucination Detection Isn't Enough — You Need a Judge, Not Just a Classifier

5 min read

A detector answers "is this output likely false?" Your risk committee asks "was this action permitted, under which rule, and can you prove it in eighteen months?" Three courts have now attributed AI statements to the operator — and none of those rulings turned on detection accuracy.

On this page

Hallucination detection is a solved-enough problem that it has become a checkbox. You can buy a classifier that scores whether an output is likely fabricated, wire it into your pipeline, and show your board a number.

Then a regulator asks a different question, and the classifier has nothing to say.

The question detection cannot answer

A detector answers: is this output likely false?

Your risk committee asks: was this action permitted, under which rule, and can you prove it eighteen months from now?

Those are not the same question, and no amount of accuracy on the first produces an answer to the second. A detection score is a probability about text. An examiner wants a decision about conduct, traceable to a written policy.

This distinction has stopped being academic.

Liability has already attached

In 2024 the British Columbia Civil Resolution Tribunal held Air Canada liable after its chatbot invented a bereavement-fare policy. The airline argued, in effect, that the chatbot was responsible for its own statements. The tribunal rejected that.

In May 2026 the Higher Regional Court of Hamm went further: a website chatbot is legally part of the business organisation, and its invented statements are attributed to the operator even where the system was carefully configured. Careful configuration was not a defence.

Weeks later, the Munich Regional Court barred Google from publishing false AI Overviews, treating the AI output as a new statement by the publisher rather than reproduced third-party content.

Three jurisdictions, one direction of travel. "The model said it, not us" has failed as a defence.

Note what none of those rulings turned on. Not detection rates. Not model accuracy. They turned on whether the organisation could be held to what its system said. The remedy a court is interested in is not a better classifier.

What a judgment is, as distinct from a score

A detector asks whether text looks fabricated. A judge asks whether an action was permitted, and answers with a reason.

The practical differences:

It reads your policy, not a general notion of truth. "Correct" in a regulated context does not mean factually accurate in the abstract. It means consistent with your exposure limits, your eligibility rules, your disclosure requirements, your procedures. Those live in documents, not in model weights.

It reasons about the whole action in context. The costly failure is rarely a false sentence. It is a consequential action produced by a defective chain of reasoning, where each individual step is well-formed. Catching that requires evaluating the action against the state of the system, not scoring a string.

It can return more than pass or fail. A judge that only ever fails things is a filter with a longer runtime. Real judgment includes the conditional case: this action is substantially correct but omits a required control, and here is the clause that requires it. It also includes recognising a correct refusal as correct — an agent declining to do something unsafe should pass, not trip a keyword rule.

It names the rule. This is the part that produces an artifact. "Blocked, 0.87 confidence" is not evidence. "Blocked under clause 4.2, exposure ceiling exceeded for this account class" is.

Why the policy corpus has to be retrieved, not memorised

A judge that has learned your policies during training is out of date the moment the policy changes — and in regulated work, it changes.

Retrieving the governing clause at the moment of judgment means a policy update is a document update. No retraining, no revalidation, no window where the judge is quietly enforcing last quarter's rules. It also means the verdict can cite the clause, because the clause was actually read.

And why it has to run where your data is

If your judge is a cloud API, then every prompt, every trace, every customer record under evaluation leaves your perimeter in order to be judged. For a bank, an insurer or a plant, that is not a procurement objection — it is the reason the project does not happen.

A judge small enough to run inside your own network, on your own hardware, removes the exchange entirely. There is no egress to negotiate because there is no egress.

The standard nobody can point to yet

The EU AI Act's high-risk provisions came into force on 2 August 2026. Autonomous systems taking consequential actions fall in scope, with obligations around human oversight, accuracy and logging.

The EU AI Office has published no agent-specific guidance. The AI Act Service Desk describes its own considerations on agents as preliminary. Independent analysis has concluded the technical standards under development are unlikely to fully address agent risks.

So compliance cannot currently be purchased as a certificate. There is no certificate. What can be produced is evidence: a per-action record showing what was attempted, what the policy required, what was allowed or blocked, and what a human did next.

That is the case a compliance officer can actually make. A detection score is not.

What to ask a vendor

Five questions that separate a judge from a classifier:

  1. Does it read our policies, and what happens when we change one?
  2. Does it evaluate actions or text?
  3. Can it return a conditional verdict, and does it correctly pass a correct refusal?
  4. Does the output name the specific rule, or produce a score?
  5. Where does it run, and what leaves our network?

Where we stand

Nucleus, our judge model, is built on open weights — we say so before anyone asks, because the base model is not the point. What matters is that it is grounded in your policy corpus retrieved at inference time, that it runs inside your network with no egress including on edge hardware, that it is independent of the agent it judges, and that every verdict names the rule it applied.

The honest boundary: judgment, blocking, escalation and isolation are live. Automatic isolation triggers on spend and cost ceilings. We do not have shipped vertical policy packs for any regulated domain, and we do not certify anyone's compliance. Domain policy packs are built with design partners, against their real documents. Certification is not a thing a vendor can sell you here — evidence is what exists, and evidence is what we produce.

The fastest way to see the difference is to take a sample of your own agent traces, run a detector over them, then run a judge over them, and compare what each one can tell your compliance officer.