New on SSRN: Ariadne's Thread, a measurement-theoretic method for legal openness. Read the paper

Good Law for What?

A Proposition-Usability Model for Verification in AI-Assisted Legal Research

Six panels, each showing the same figure in a dark coat and cap from behind, facing a different hall of authority rendered through a haze of data; two panels reveal a domed capitol building.
Plate 27 · Six CorridorsPlates
Abstract

Verification as currently taught asks whether a case is good law. Attorneys need to know whether a proposition is usable: still valid, governing in the relevant forum and procedural posture, helpful to the represented party, and not already rejected in the jurisdiction. Each concept exists somewhere in research competencies or commercial tools, but they remain dispersed, never unified into a proposition-level verification protocol for AI-assisted legal research. This article proposes that unification as a proposition-usability model, assesses how far Shepard's, KeyCite, and Westlaw Precision already reach, and offers a classroom exercise and vendor-evaluation questions law librarians can apply now.

Authors
Ross Brodskiy and Nathan Pokov
Posted
1 September 2026 (SSRN)
SSRN abstract
7373039
Length
40 pages
DOI
Not yet assigned
Keywords
proposition usability, verification, citators, AI-assisted legal research, law librarianship, legal research instruction
Key findings
  1. Verification as ordinarily practiced collapses four questions into one. It asks whether a case is good law, when the attorney needs to know whether a proposition is still valid, governs in this forum and posture, helps the client, and has not already been rejected here.
  2. The hardest changes are not treatment events. After Kisor v. Wilkie, Auer deference survived but its practical direction shifted toward challengers, and no citator vocabulary has a category to record that.
  3. Shepard's headnote restrictions, KeyCite Overruling Risk and Overruled in Part, and Westlaw Precision's motion, party and outcome classifications each reach part of the protocol; none assembles all four questions.
  4. A flawless citator would still leave the gap open, because the limitation is representational rather than editorial: transplantability, direction and refused contentions have no place in the data model.
  5. The distinction between structural and opportunistic reliability turns an unanswerable question (is this tool accurate?) into an answerable one (for which classes of error does it have a defined detection procedure?).
  6. Law librarians can adopt the model now: four teaching questions, a classroom exercise with an assessment rubric, and seven demonstration-based procurement questions that the author's own products can fail.
Contents of the paper
  • I. Introduction: The Visible Failure and the Invisible One
  • II. What Verification Currently Means
  • III. The Proposition-Usability Model: Four Verification Questions
  • IV. What the Commercial Tools Already Do, and Where They Stop
  • V. Why Better Editorial Performance Would Not Close the Gap
  • VI. What Generative Research Tools Change
  • VII. Implications for Law Librarianship
  • VIII. Conclusion
  • Appendix: A Classroom Exercise and Assessment Rubric

The paper in brief

Since Mata v. Avianca, the profession has taken one lesson from artificial intelligence in legal research: check that every cited case exists and says what the brief claims. Brodskiy and Pokov call this the right response to the failure we can see and an incomplete response to the failures we cannot. A fabricated citation announces itself. The errors this paper cares about involve real cases, accurately quoted and correctly Shepardized, cited for propositions they will not bear in this jurisdiction, at this procedural stage, or for this party. Nothing flags them.

Part I: the visible failure and the invisible one

The cause is structural and predates generative AI. The verification apparatus answers a question about cases; attorneys need an answer about propositions. A case can live while the proposition drawn from it is dead, displaced in this forum, pointed at the wrong client, or already refused here. For a century the profession closed that gap by having a lawyer read the case. Generative tools do not create the gap. They remove the friction that used to expose it, producing fluent memoranda with real citations that a citator will bless.

The paper is careful about novelty. Each of the four attributes it names already lives somewhere in the profession's dispersed apparatus: headnote-restricted citator reports, jurisdiction filters, case analysis, opposition research. The claimed contribution is the assembly: one named protocol that asks all four questions about every load-bearing proposition, can be taught in one class session, can be demonstrated against any product without vendor cooperation, and is stated independently of any system's architecture.

Part II: what verification currently means

Two nineteenth-century inventions still carry the load. Shepard's lists (from 1873) track what later courts did to a case; West's Key Number System breaks opinions into editorially written points of law. The first is a relationship-level artifact, the second a proposition-level one. Drawing on Berring and on Callister, the paper notes that what the editorial system declines to represent tends to become invisible, and that generative systems now speak the law itself, so authority migrates into the answer.

A red flag says some later authority did something significantly negative to the case. It does not say the quoted passage was affected, or that the treatment binds in this forum. The empirical record is not reassuring. Kirschenfeld found treatment vocabularies applied inconsistently. Hellyer found that Shepard's and KeyCite each missed or mislabeled about a third of negative citing relationships, and that all three services he tested agreed on negative treatment in only 53 of 357 relationships. Mart found that the same search run in six databases produced top-ten results that were, on average, 40 percent unique to one database.

Part III: four verification questions

A case is not a unit of law. The attorney cites a proposition and offers a case as its container. A proposition is usable when it satisfies four questions.

  1. Is it still good law, as distinct from the case? Loper Bright overruling Chevron is the easy test. Kisor v. Wilkie is the hard one: it preserved Auer deference but conditioned it on gates that each open a door for the challenger. The citator entry is accurate and records nothing, because a change in a doctrine's practical direction is not a treatment event.
  2. Does it govern in this forum and posture? Federal summary judgment under Celotex lets a moving defendant point to an absence of evidence; California under Aguilar requires more. The two lines share vocabulary and appear together in any competent search. The tell worth teaching is a persuasive citation from another system doing the work of a controlling one.
  3. Whom does it help? HYPO and CATO treated direction as foundational decades ago. In today's platforms, directional stance is still not a first-class attribute of a point of law.
  4. Has it already been rejected here? A refused contention is rarely a headnote and rarely produces a citator signal. Finding it requires a search about the argument you are about to make. This is the profession's negative space.

Part IV: what the commercial tools already do

The paper gives the incumbents full credit. Shepard's can restrict citing references to a headnote and attach positive and negative labels to the same citing decision. KeyCite Overruling Risk (2018) infers implicit undermining at the paragraph level, and Overruled in Part (2022) addresses partial invalidity. Westlaw Precision, launched in 2022 with more than 250 newly hired attorney editors, classifies covered cases by issue, outcome, motion type and party type. It is the closest approach, but its coverage is bounded, its attributes serve retrieval rather than verification, and a particular sentence can fall between its tags. Efforts to rebuild the citator in the open, or to have courts produce their own signals, concern who makes treatment signals; any treatment signal answers at most the first question.

Part V: why better editorial performance would not close the gap

Imagine a citator that never misses a negative event and a classification system with perfect recall. The citator would still miss the Kisor class and say nothing about transplantability, direction or refused contentions. The limitation is representational, not editorial: the data model was settled when every proposition passed through a reading lawyer.

Part VI: what generative tools change

Studies show retrieval-augmented tools reduce hallucination without eliminating it, but the fabrication metric is orthogonal to the four questions. The paper then defines two kinds of reliability. A tool is structurally reliable for a class of error when the information needed to catch it is recorded in the tool's data and consulted by a defined, repeatable check. It is opportunistically reliable when correctness depends on whether retrieval happens to carry the warning. Loper Bright is a weak test of any tool; Kisor is the strong one.

Parts VII and VIII: implications and conclusion

For teaching, "verify the case" becomes four questions asked about each proposition. For procurement, seven questions are answered by demonstration, not assertion: whether compound constraints are honored, whether validity is tracked below the case, whether shared-vocabulary doctrines stay apart, whether rejected arguments are searchable, whether the reasoning trace is inspectable, whether coverage and update latency are disclosed, and whether classifications carry human editorial provenance. The set is rigged against no product class; incumbents tend to pass the last two, newer systems the first four. On competence, ABA Formal Opinion 512 is right but does not say what verification consists of, so the paper proposes a disclosure norm: products should state which classes of error they are designed to detect.

The appendix supplies a ready-to-use exercise. Students receive a machine-drafted memorandum on the movant's initial summary judgment burden in California Superior Court, list its propositions (not its cases), and complete a four-row verification table in which every answer rests on a citation or a recorded search. Split the class between movant and opponent, and question three returns different answers about the same propositions, which is the point.

Disclosure. The paper states that Ross Brodskiy has a financial and proprietary interest in commercial software that performs proposition-level legal knowledge extraction and citation validation, with related patent applications pending, and that this interest is why it proposes a standard his own products can fail, stated so any librarian can apply it without vendor cooperation, and describes no proprietary system.

Where to go next

Frameworks in this piece

Terms in this piece

Revision history

20 Sep 2026Paper page created.

How to cite

Brodskiy, R., & Pokov, N. (2026, September 1). Good Law for What?: A Proposition-Usability Model for Verification in AI-Assisted Legal Research [Working paper]. SSRN. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7373039

Related pieces

Follow the research

New papers, frameworks and essays. No marketing. Or use RSS.