New on SSRN: Ariadne's Thread, a measurement-theoretic method for legal openness. Read the paper
10 problems

Open Problems

What we do not yet know how to do. Each problem states why it matters, the state of the art, our partial answers and an honest status.

openMeasuring Determinacy Across JurisdictionsHow can the openness of a legal question be measured comparably across forums whose procedural rules, pleading standards and court structures differ?DeterminacyopenPosture SeparabilityCan language-model reasoners hold genuinely distinct judicial postures over the same correct law, or do they collapse toward a single default stance?Determinacypartially addressedBenchmarking Beyond Citator FlagsWhat public, rerunnable benchmark can measure whether a legal research system finds the controlling authority and a proposition usable in the forum, rather than only whether its citations exist and carry a green flag?VerificationopenDetecting Omissions from the Document AloneCan missing controlling or adverse authority be detected by examining a finished brief, without access to the research process that produced it?VerificationopenState-Court Pleading Coverage in Training DataHow far does the dominance of federal practice in written legal material bias AI systems toward federal standards in state-court work, and how can that bias be measured and corrected?VerificationopenTreatment Label Disagreement Among CitatorsWhen citators disagree about how a later case treated an earlier one, what should count as ground truth, and how should a system act on a contested label?Knowledge EngineeringopenAuditable Multi-Agent Legal WorkHow can legal work divided among several AI agents be made auditable enough for a lawyer to sign, with a record of which agent did what, which decisions were reserved to humans, and why the system stopped?System Designpartially addressedCalibrating Confidence in Legal OutputsWhen should a legal AI system express confidence as a calibrated number, when in words, and when not at all?System DesignopenImplicit Responsiveness in eDiscoveryHow can automated review find documents that are responsive without saying so, such as chats, meeting minutes and presentations whose relevance is implicit?System Designpartially addressedA Public Record of Citation FailuresHow should AI citation failures in court be recorded and classified so that bar associations, judges and malpractice carriers can set policy on evidence rather than anecdote?Machine Jurisprudence