New on SSRN: Ariadne's Thread, a measurement-theoretic method for legal openness. Read the paper

Measuring Doctrinal Indeterminacy

Why legal AI must distinguish settled law from contested doctrine, and how the distinction can be operationalized

Two luminous blue waveforms spiral around each other across a dark field and converge at a single bright point of white light, like a gravitational-wave signal emerging from noise.
Plate 35 · The WaveformPlates

Discussion of the reliability of legal AI tends to conflate three distinct questions. The first is a question of integrity: does the cited authority exist, and does it say what the system claims? The second is a question of completeness: has the system found the authority that exists? The third, and least examined, is a question of determinacy: does the legal system itself treat the question as settled?

The literature on "hallucination" addresses the first question, and retrieval benchmarks address the second. Neither addresses the third. Yet a system can pass both of the first two tests and still misstate the law. It may retrieve five cases supporting one position and five supporting the opposite, then synthesize them into a confident statement that the law is the first position. Nothing has been fabricated; nothing relevant has been missed. The system has simply presented contested doctrine as settled.

Call this failure the indeterminacy gap: a system's inability to distinguish I do not know the answer from there is no determinate answer. The first state calls for better retrieval. The second calls for a flag. What follows moves from definition to taxonomy, architecture, measurement, and objections.

I. Two kinds of uncertainty

Two terms of art organize the argument.

Shallow ambiguity arises from incomplete information. The resolution exists in the legal materials but has not been located. A question of first impression in one jurisdiction that other jurisdictions have resolved is shallowly ambiguous in this sense. Shallow ambiguity is a property of the research process.

Genuine ambiguity arises from the structure of the law itself. No additional research resolves it, because the materials support incompatible readings or the system treats the question as contested, as in a persistent circuit split. Genuine ambiguity is a property of the legal system.

The two call for different professional conduct. Shallow ambiguity is remedied by research. Genuine ambiguity is not remedied at all; it is disclosed, briefed on both sides, and priced into advice.

For purposes of measurement, the Institute's working drafts adopt the following definition:

Genuine legal indeterminacy exists when the legal system itself treats a doctrinal question as contested: that is, when courts of competent jurisdiction have produced holdings in good faith that support materially different outcomes, each grounded in legitimate interpretive methodology.

The definition has three features. It is observable: divergent holdings are a fact about the corpus. It is systemic rather than epistemic: it does not depend on what any researcher knows. And it is threshold-based: a lone outlier does not make settled law indeterminate, and a disagreement a higher court has resolved no longer signals indeterminacy. What counts is sustained, reasoned divergence among courts of equal authority.

II. Why an operational definition

The operational definition deliberately brackets a metaphysical dispute.

H.L.A. Hart argued that legal rules have a core of settled meaning and a penumbra of uncertain application.1 A rule prohibiting vehicles in a park is settled for automobiles and uncertain for bicycles, roller skates, and skateboards. The penumbra is not a drafting defect; it follows from applying finite general language to an unbounded variety of cases. But Hart offered no systematic method for telling which applications fall in the core and which in the penumbra. That is precisely what a computational system must supply.

Ronald Dworkin contested the inference from hard cases to indeterminacy.2 Even penumbral cases, on his account, have right answers, which an idealized judge, Hercules, would find in the interpretation that best fits and justifies the law. A pragmatic reading of Dworkin nonetheless concedes what matters here: right answers may exist in principle yet be practically inaccessible to real judges, lawyers, and systems under constraint. While competent interpreters disagree, courts may reach opposite conclusions, and neither should be presented as settled.

Karl Llewellyn's observation that the canons of construction come in opposed pairs locates further uncertainty in the interpretive tools themselves.3 Critical Legal Studies generalized the point into a thesis of pervasive indeterminacy. Its most useful contribution, for present purposes, is identifying where indeterminacy concentrates: multi-factor balancing tests, open-ended standards, and doctrine that depends on policy trade-offs.

The operational definition is neutral among these positions. It asks not whether a question has a right answer but whether the legal system presently treats it as contested. A positivist and an interpretivist can agree on that observation while disagreeing about its significance. Neutrality of this kind is not evasion; it is what makes the construct measurable.

III. A taxonomy of contestation

Genuine indeterminacy has at least five sources. Each leaves a different trace in the corpus and requires a different detection strategy; a system that treats all uncertainty alike will generate false positives.

TypeTheoretical sourceWhat the corpus showsDetection feasibility
SemanticHart's open textureDivergent judicial constructions of the same termModerate
NormativeDworkin's hard cases; Alexy's balancingShared principles weighted differently; unranked multi-factor testsModerate to hard
MethodologicalLlewellyn's dueling canons; Eskridge; GroveMajority and dissent invoking opposed interpretive methodsHard
JurisdictionalCircuit splitsCourts of equal authority holding opposite rulesHigh
AnalogicalLevi; SunsteinDifferent precedents treated as controlling for similar factsModerate to hard
  1. Semantic indeterminacy: a legal term admits several defensible interpretations, none clearly primary. Miller v. California left "patently offensive" undefined,4 and Justice Stewart's concurrence in Jacobellis v. Ohio conceded that he might never define the category intelligibly: "But I know it when I see it."5 The signal is a single predicate applied to different entities with different outcomes.

  2. Normative indeterminacy: principles of roughly equal weight pull in different directions, and the balance is not dictated by the principles themselves. New York Times Co. v. Sullivan balanced the First Amendment against the protection of reputation and produced a standard that privileges speech without eliminating reputational protection.6 The markers are "balance," "weigh," and "on balance." Detection compares the normative profiles of holdings that cite the same materials but rank them differently; ground truth is necessarily fuzzy.

  3. Methodological indeterminacy: legitimate interpretive methods yield different results from the same materials. The working drafts take the overruling of Chevron U.S.A., Inc. v. Natural Resources Defense Council, Inc.7 in Loper Bright Enterprises v. Raimondo8 as their paradigm, reading majority and dissent as applying opposed methods to the same question. This is the hardest type to detect reliably, though methodological disagreement is often explicit in the opinions.

  4. Jurisdictional indeterminacy: courts of equal authority hold materially different rules. Jurisdiction is explicit metadata, so detection is comparatively easy. The difficulty is definitional: whether minor doctrinal differences count, or only different outcomes. A split is also forum-relative; a question may be settled within each circuit and indeterminate only nationally.

  5. Analogical indeterminacy: several legitimate precedents support different outcomes by analogy, and no hierarchy among them resolves the conflict. AI-generated creative content can be analogized to photography, to a human using a tool, or to curated found material, and each characterization yields a different conclusion about protection.

Note the inverse relation between familiarity and difficulty. The type practitioners know best, the circuit split, is the easiest to detect; the types that most trouble theorists are the hardest.

IV. Why retrieval cannot see contestation

Retrieval-augmented generation has genuine strengths: it grounds a model in actual documents, scales well, and reduces fabrication. Its limitation here is structural, in three respects.

The unit of retrieval. Retrieval operates on text chunks; doctrine is organized as propositions and their relationships. To ask whether a jurisdiction recognizes strict liability for abnormally dangerous activities is to ask how the holdings on that question are distributed, a question about the corpus as a whole.

Metadata as an afterthought. Retrieval can be supplemented with court, date, and jurisdiction, but its core operation is semantic similarity. Contestation is defined almost entirely by the metadata similarity search de-emphasizes: who decided, when, with what authority, and whether one court disagrees with another or merely distinguishes it.

Relevance is not stance. Two passages may be equally relevant and hold opposite positions. A model optimized for fluent synthesis will tend to reconcile them rather than report their conflict.

Existing tools do not fill the gap. Citators classify treatment, not the aggregate status of a doctrine; a case may be distinguished many times without any flag that the doctrine is contested. Topic systems sort by subject, not by agreement. CaseHOLD, with more than 53,000 annotated holdings,9 shows that holdings can be extracted at scale, but an isolated holding says nothing about whether its doctrine is settled. Overruling detection answers a binary, per-case question; indeterminacy detection answers an aggregate one.

The drafts propose instead a knowledge graph whose nodes are holdings carrying metadata (stance, origin, court, date, holding or dicta, precedential weight) and whose edges record typed relations: elaboration, application, narrowing, overruling, unresolved tension, analogy. Detecting contestation then becomes a graph operation: gather the holdings on a proposition, cluster them by stance, assess the configuration by jurisdiction and authority, and propagate the result to dependent doctrines.

One commitment answers an obvious worry. Language models are used for extraction and classification, not for the determination itself. Whether a doctrine is contested should follow from the structure of the holdings, not from a model's judgment. A model's uncertainty about its output and the law's uncertainty about a question are different quantities.

V. From signals to a measure

Evidence of contestation falls into four families: explicit (courts saying so: "the circuits are divided," "first impression"); structural (framework-level dissents, narrowing holdings, unranked balancing tests); aggregate (splits, and one principle appearing in holdings of opposite stance); and temporal (drift, emerging splits, splits later resolved).

The Institute's paper Detecting Genuine Doctrinal Ambiguity organizes these into three layers. Layer A reads individual opinions at extraction time (split acknowledgments, hedging, dissent vigor, canon conflict, and similar signals). Layer B aggregates across the corpus (stance inversion, jurisdictional divergence, drift, outcome variance, dissent and reversal rates). Layer C evaluates doctrinal structure (fragmentation into inconsistent sub-rules, courts applying different tests to the same question). Eighteen tests feed a weighted composite, the Grayness Score, reported from 0 to 1 in tiers running from clear law to profound indeterminacy.

Grading rather than classifying is itself a substantive claim. A 5-4 split among circuits is plainly contested; an 8-1 split arguably so; universal circuit agreement with one dissenting district court probably not. A binary detector would discard exactly the information a practitioner needs.

Equally important is what accompanies the number: the evidence that produced it. An attorney can open the cited opinions and verify that the conflict exists. The explanation, not the score, distinguishes this construct from a model's confidence estimate.

The paper's worked illustration is Collins v. Virginia,10 which sits where the automobile exception meets the protection of curtilage. The majority reserved questions about other settings; the dissent predicted the rule would prove unworkable; lower courts, on the paper's account, have since diverged on its reach. The illustration yields a score in the significant-gray-area tier. It is expressly illustrative, not an evaluated result.

Validation

How does one validate a system whose correct output, in hard cases, is that there is no correct answer? The paper constructs ground truth from three sources (documented splits and certiorari grants, expert ratings on a five-point scale, and outcome signals such as reversal rates) and fixes success and failure conditions in advance. Success requires a Spearman's ρ above 0.7 against expert judgment, with inter-rater reliability at a Cohen's κ above 0.6. The framework counts as failed if the correlation falls below 0.3, or if experts find more than 20 percent of its cited conflicts spurious. It must also outperform a keyword heuristic, a citation-network measure, and direct prompting of a language model. Testbeds stress different types: employment discrimination (normative), arbitration (methodological), and the Fourth Amendment (jurisdictional).

Pre-registered failure conditions are the feature most worth defending. They convert a design proposal into a falsifiable claim.

VI. Objections

  1. The interpretivist objection. If hard cases have right answers, a grayness score measures disagreement, not indeterminacy. The objection is correct about the construct and harmless to the project. What is measured is institutional contestation, and it would be more precise to call it that. A lawyer's duties to disclose, to brief both sides, and to advise on risk are triggered by contestation, whatever its ultimate ground.

  2. The ground-truth objection. Expert ratings are themselves contestable, so validation against them is circular. The reply is partial: expert disagreement is measured, and ratings are triangulated against formal acknowledgments and outcome signals. Circularity is reduced, not eliminated, and a published validation should say so.

  3. The classification objection. The five types are not mutually exclusive. One working draft files Collins under semantic indeterminacy; the paper reads it as a collision of two doctrinal lines, closer to normative or analogical contestation. The better conclusion is that the types are sources, not bins. A score should record which sources contribute to it rather than force a question into one type.

  4. The counting objection. One draft proposes a rule of thumb: if 80 percent of holdings favor one position, the doctrine is settled though contested; an even division signals clear indeterminacy. Raw ratios are too crude. Holdings differ in authority, and the relevant population is forum-relative: a question may be settled where the client litigates and open elsewhere. An aggregate measure must weight by authority and report determinacy relative to a forum.

  5. The temporal objection. A system can identify a doctrine's current status; it cannot predict whether that status will hold. This objection admits no technical answer. It marks the boundary of the project: the measure describes the present landscape, and judgment about where it is moving remains human.

A practical concern remains. Knowledge acquisition has always been computational law's bottleneck, and a graph with stance and typed relations is expensive to build. The drafts concede that comprehensive federal coverage would be substantial and a narrower domain manageable; early claims should be domain-bounded accordingly.

VII. Professional stakes

The duties engaged are familiar. Model Rule 1.1 requires competence; Rule 1.4 requires keeping clients reasonably informed; Rule 3.3 requires disclosure of directly adverse authority. The last is narrower than it is sometimes described: on its terms it reaches adverse authority in the controlling jurisdiction. A split confined to other courts engages it less directly than it engages competence and communication, because advice that presents a contested question as settled misinforms the client about risk. For self-represented litigants, with no counsel to supply the missing caveat, overconfident guidance falls hardest.

VIII. Conclusion

The indeterminacy gap will not be closed by making law determinate, which is impossible. It can be narrowed by making the law's indeterminacy explicit. That requires systems that:

  1. represent holdings as propositions, not text chunks;
  2. carry metadata about stance, jurisdiction, date, and weight of authority;
  3. detect contestation through explicit, structural, aggregate, and temporal signals;
  4. report a graded, forum-relative measure together with the evidence behind it;
  5. distinguish "I have not found it" from "the courts do not agree"; and
  6. present themselves as augmenting, not replacing, legal judgment.

This is not skepticism about legal AI. It is a claim about what such systems owe their users: an honest account of what the law settles, what it leaves open, and how to tell the difference. Hart described the core and the penumbra. The task now is to measure where one ends and the other begins, and to show the evidence.

Footnotes

  1. H.L.A. Hart, The Concept of Law (1961; 2d ed. 1994). ↩

  2. Ronald Dworkin, Taking Rights Seriously (1977); Ronald Dworkin, Law's Empire (1986). ↩

  3. Karl N. Llewellyn, "Remarks on the Theory of Appellate Decision and the Rules or Canons About How Statutes Are to Be Construed," 3 Vanderbilt Law Review 395 (1950). ↩

  4. Miller v. California, 413 U.S. 15 (1973). ↩

  5. Jacobellis v. Ohio, 378 U.S. 184, 197 (1964) (Stewart, J., concurring). ↩

  6. New York Times Co. v. Sullivan, 376 U.S. 254 (1964). ↩

  7. Chevron U.S.A., Inc. v. Natural Resources Defense Council, Inc., 467 U.S. 837 (1984). ↩

  8. Loper Bright Enterprises v. Raimondo, 144 S. Ct. 2244 (2024). ↩

  9. L. Zheng et al., "When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset of 53,000+ Legal Holdings," Proceedings of the 18th International Conference on Artificial Intelligence and Law (2021). ↩

  10. Collins v. Virginia, 138 S. Ct. 1663 (2018). ↩

Frameworks in this piece

Terms in this piece

Revision history

23 Jul 2026Rewritten for the Institute library by Priya Raman.

How to cite

Raman, P. (2026, July 23). Measuring Doctrinal Indeterminacy: Why legal AI must distinguish settled law from contested doctrine, and how the distinction can be operationalized. Computational Law Institute. https://institute.legawrite.ai/articles/measuring-doctrinal-indeterminacy

Related pieces

Follow the research

New papers, frameworks and essays. No marketing. Or use RSS.