New on SSRN: Ariadne's Thread, a measurement-theoretic method for legal openness. Read the paper
23
Legal AI System Design

A benchmark that grades only what is present grades the wrong thing.

First stated 20 September 2026 in The Verification Layer Is a Smoke Alarm, Not a Building Code

Defense

Most legal AI comparisons ask two questions: did the system find a relevant case, and did it invent a citation? Grounded systems now pass both. A system that retrieves from a real corpus and checks against primary sources can report a fabrication rate of zero. If the benchmark stops there, it cannot tell a filing-grade tool from one that returns real, perfectly formatted cases that are wrong for the job.

The failures that decide motions are absences and mismatches: the controlling case never surfaced, the adverse authority never shown, the element never addressed, the holding from the wrong stage, the case reversed last month. A benchmark that scores only the contents of an answer inherits the verifier's blind spot. In our survey of public legal AI benchmarks, none measured failure to retrieve controlling authority.

A filing-grade benchmark measures dispositive recall under constraint: whether the system surfaced the proposition that does the legal work, with the correct stance, at the correct stage, still good law, and currently citable in the forum. The constraints compose multiplicatively. A system can clear any one of them by luck; far fewer clear all of them together. Such a benchmark also keeps the topical-recall row in, because showing where competitors are strong is what makes a gap elsewhere credible.

Two reporting rules keep it honest. Do not average per-class results into one score, because a tool that is perfect on fabrication and poor on controlling authority will look trustworthy while missing the case that matters. And publish the rubric so that anyone can rerun it. A scoreboard nobody can rerun is marketing.

Supporting pieces

Related framework

Revision history

20 Sep 2026Added to the Theses.

How to cite this thesis

Computational Law Institute (2026, September 20). Thesis 23: A benchmark that grades only what is present grades the wrong thing.. https://institute.legawrite.ai/agenda/theses/23

Cite the thesis by number and the date of the revision you relied on.