New on SSRN: Ariadne's Thread, a measurement-theoretic method for legal openness. Read the paper
partially addressedVerification and Usable Law

Benchmarking Beyond Citator Flags

What public, rerunnable benchmark can measure whether a legal research system finds the controlling authority and a proposition usable in the forum, rather than only whether its citations exist and carry a green flag?

Why it matters

Grounded systems can now report fabrication rates of zero. The evaluation vocabulary has not caught up. "Our citations are verified" answers a question about existence and current flag, not about whether the proposition governs in the forum, at the stage, for the client, or whether the controlling case was found at all. Without a benchmark that measures those things, procurement, bar guidance and malpractice underwriting rest on vendor assertions.

State of the art

Peer-reviewed testing of commercial retrieval-augmented research tools (Magesh et al., 2025) labeled answers as grounded, ungrounded, misgrounded or fabricated and found hallucination rates between 17 and 33 percent, but did not score controlling authority that was never retrieved. Broad reasoning suites such as LegalBench do not measure citation accuracy or omission. Vendor metrics of groundedness ask whether an answer is supported by some source, not by the right one. In our survey, no public benchmark measured failure to retrieve controlling authority. Citator research adds a further complication: the flags themselves are contested.

Our partial answers

Three instruments cover parts of the problem. The Docket Test uses a decided motion, with its opposition and order, as adversarially produced ground truth, and scores five lines side by side without averaging: controlling authority omitted, adverse authority surfaced, posture mismatches, element coverage, and fabricated or misgrounded citations. The proposition-usability model supplies four verification questions and a classroom rubric that a librarian can apply without vendor cooperation. The filing-grade benchmark design measures dispositive recall under multiplicative constraints (right proposition, stance, stage, good law, currently citable), keeps the topical-recall row in, and requires the rubric to be published and rerunnable.

Still open: a shared public corpus with answer keys across several jurisdictions; independent administration, so that no vendor grades itself; and a principled way to handle orders decided on narrow grounds and briefs that omitted authority for strategic reasons.

Linked pieces