New on SSRN: Ariadne's Thread, a measurement-theoretic method for legal openness. Read the paper
partially addressedLegal AI System Design

Calibrating Confidence in Legal Outputs

When should a legal AI system express confidence as a calibrated number, when in words, and when not at all?

Why it matters

The Zeroth Law forbids presenting output with more confidence than the system can defend. That requires knowing how much confidence can be defended, and expressing it in a form a lawyer will read correctly. An uncalibrated confidence score is worse than none, because it invites reliance. Yet a system that never quantifies its confidence cannot be tested on it.

State of the art

Classification tasks with labeled data can be calibrated with standard methods such as Platt scaling or isotonic regression, and calibration can be measured, for example by expected calibration error. Most legal research outputs are not labeled classifications. Whether a distinction holds or a proposition governs is a legal judgment with no ground truth at the moment of use. Confidence displays in research tools often reflect retrieval similarity or citation frequency rather than doctrinal uncertainty.

Our partial answers

Our work has reached two different answers for two kinds of output, and reconciling them is part of the problem. In technology-assisted review, where responsiveness can be labeled and sampled, every decision carries a calibrated confidence and calibration is a gate. In the blind eDiscovery review, expected calibration error rose to 0.088 against 0.046 in the previous batch; the gate failed, automatic acceptance was locked, and all 2,038 responsive determinations went to human review. In opposition drafting, where the question is whether a legal distinction holds, the design reports strength in words (strong, moderate or weak) by category and never as a number, on the view that an 87 percent score on a distinction is an invitation to stop thinking. Ariadne's Thread reports openness only as an interval, with abstention.

Open: a principled rule for which outputs can carry a calibrated number; a way to calibrate confidence in legal propositions without labels; and a test of whether verbal strength categories are applied consistently.

Linked pieces