New on SSRN: Ariadne's Thread, a measurement-theoretic method for legal openness. Read the paper

Detecting Genuine Doctrinal Ambiguity

A Multi-Layer Framework for Identifying Structural Indeterminacy in Judicial Reasoning

A figure in a dark suit stands on a rocky ledge beside a floating display, overlooking a storm-dark mountain valley laced with glowing green paths and crossed by a single red band.
Plate 25 · Fault LinesPlates
Abstract

Current AI systems suffer from 'false confidence,' presenting contested legal conclusions as settled. While computational law has modeled reasoning, a gap exists in detecting where reasoning breaks down due to structural indeterminacy. This article addresses this by distinguishing between 'shallow ambiguity' (resolvable by research) and 'genuine ambiguity' (structural features of law). We propose a five-type taxonomy (semantic, normative, methodological, jurisdictional, and analogical) accompanied by the Gray Area Detection Framework. This three-layer architecture utilizes extraction-time signals, query-time aggregation, and structural analysis to identify relational ambiguity across large corpora. By shifting focus from outcome prediction to contestation detection, the framework mitigates LLM-era risks like hallucination and decontextualization. We specify a validation methodology using expert annotation, formal acknowledgments, and domain-specific testbeds (Employment Discrimination, Arbitration, Fourth Amendment) to measure system utility. Ultimately, this approach moves legal AI from an 'oracle' model toward a systematic 'cartographer' of the law's gray areas, enhancing judicial transparency and reinforcing the rule of law through verifiable explanations of doctrinal conflict.

Authors
Ross Brodskiy
Posted
23 July 2026 (SSRN)
SSRN abstract
7039398
Length
13 pages
DOI
Not yet assigned
Keywords
legal indeterminacy, computational jurisprudence, doctrinal conflict, gray area detection, legal AI safety, case-based reasoning
Key findings
  1. Shallow ambiguity is a research problem that more searching resolves; genuine ambiguity is a structural feature of the law that no amount of research removes.
  2. An operational definition sidesteps the Hart and Dworkin debate: a question is genuinely ambiguous when the legal system itself treats it as contested, through splits, hedging or vigorous dissent.
  3. Genuine indeterminacy comes in five detectable types (semantic, normative, methodological, jurisdictional and analogical), each tied to a theoretical source and a distinct detection strategy.
  4. Ambiguity is relational: a circuit split is invisible in any single opinion. Detection therefore needs three layers (extraction-time signals, query-time aggregation, structural analysis), implemented as 18 tests feeding a composite Gray Area Score.
  5. Every score is grounded in citable judicial language, so an attorney can check the conflicts the system reports instead of trusting a model's confidence.
  6. The framework is built to be falsified: it must reach a Spearman correlation above 0.7 with expert judgment, beat keyword, citation-network and direct-prompting baselines, and is treated as failed if correlation falls below 0.3 or more than 20 percent of cited conflicts prove spurious.
Contents of the paper
  • Introduction: The Problem of Uncharted Gray Areas
  • Jurisprudential Foundations: Defining Genuine Ambiguity
  • Why Existing Computational Approaches Fail to Detect Indeterminacy
  • The Multi-Layer Detection Architecture
  • Validation Strategy and Falsifiability
  • Conclusion: Toward a Systematic Jurisprudence

The paper in brief

Legal AI has spent most of its effort on getting answers right. This paper argues that the more urgent missing capability is knowing when there is no single right answer to get. When a system predicts one outcome for a question that has divided the federal circuits for decades, it is not being accurate. It is displaying false confidence, and in doing so it quietly decides a contested question in a black box. Brodskiy proposes to shift the goal from prediction to contestation detection: finding the places where the legal materials themselves fail to yield a unique outcome.

Introduction: the problem of uncharted gray areas

The paper calls indeterminacy detection the "missing layer" of the legal AI safety stack. Treating unsettled law as settled is more than a technical error. It fails to respect the dialectical character of common-law reasoning and erodes the transparency on which judicial legitimacy depends. The paper makes four contributions: a distinction between shallow and genuine ambiguity, a five-type taxonomy of indeterminacy, a three-layer detection architecture, and a validation program designed so the framework can fail.

Jurisprudential foundations

The central conceptual move is the shallow and genuine distinction. Shallow ambiguity comes from incomplete information: a question of first impression in one jurisdiction that other jurisdictions have already resolved is only shallowly open, because the answer exists and has not yet been found. Genuine ambiguity comes from the structure of law itself: the materials support incompatible readings, or the system treats the question as contested, as in a persistent circuit split. More research cannot resolve it.

The paper then turns legal philosophy into checkable requirements with a five-type taxonomy.

TypeTheoretical sourceWhat it looks like in opinions
SemanticHart's open textureCourts construe the same term in divergent ways
NormativeDworkin's hard cases, Alexy's balancingCompeting principles resolved "on balance" or by unweighted multi-factor tests
MethodologicalLlewellyn's dueling canons, Eskridge, GroveMajority and dissent invoke opposing interpretive canons
JurisdictionalFederal circuit splitsCourts expressly acknowledge that the circuits are divided
AnalogicalSunstein, LeviCourts treat different precedents as controlling on similar facts

A reader steeped in jurisprudence will ask whether the project assumes Dworkin is wrong. The paper's answer is an operational stance. Even if Hercules could find a right answer in every hard case, real legal systems work under constraints that make such answers practically inaccessible. Defining genuine ambiguity as whatever the legal system itself treats as contested keeps the criterion neutral between positivism and interpretivism while remaining computable. A longer companion treatment of the same framework adds that categories can overlap (normative and methodological indeterminacy can blur), and that a question triggering several detectors is signaling compound indeterminacy, not exposing a flaw in the taxonomy.

Why existing computational approaches fail

Each earlier tradition models law on the assumption that its inputs are settled. The symbolic tradition of Allen and McCarty formalizes rules precisely, but it requires someone to resolve ambiguity before the logic can be written. Case-based reasoning in HYPO and CATO shows how to argue within a doctrine; it cannot tell when the doctrine itself is fragmenting. Argumentation frameworks in the Prakken and Sartor line formalize how rules defeat one another but presuppose that priorities between them can be specified. Legal NLP can label rhetorical roles, while large language models add hallucination and decontextualization, which the paper reads as indeterminacy failures in disguise: a model resolving a genuine tension in its training data with one confident, false statement. The summary line is that existing systems model law, and none of them detects where law runs out.

The multi-layer detection architecture

Ambiguity is rarely a property of a single document. It lives in the space between cases, and a circuit split appears only when opinions are mapped against their peers. The architecture therefore runs in three layers.

  • Layer A, extraction-time signals (8 tests), read individual opinions: explicit split acknowledgment, first-impression declarations, judicial hedging, dueling authorities, dissent vigor, multi-factor balancing, canon conflict and precedent instability.
  • Layer B, aggregate tests (6 tests), look across the corpus: stance inversion, jurisdictional divergence, temporal drift, outcome variance, dissent rate and reversal rate.
  • Layer C, structural tests (4 tests), examine the doctrine itself: fragmentation, test-standard mismatch, constitutional penumbra, and the composite Gray Area Score that normalizes and weights every contributing signal.

The ordering answers a practical constraint. Real systems ingest opinions one at a time, so single-case signals are captured at extraction and confirmed by aggregation at query time. The companion treatment sets out score bands running from clear law through significant gray area to profound indeterminacy, and explains why the design prefers missed signals to false positives at extraction.

The paper's worked illustration is Collins v. Virginia (2018), on whether police need a warrant to search a vehicle parked in a home's driveway. Layer A picks up the majority's express reservation of questions it did not decide and a vigorous dissent predicting that the rule will prove unworkable. Layer B looks for divergence among lower courts applying the decision, and Layer C for the curtilage concept splintering into sub-rules. The illustrative result is a Gray Area Score of about 0.65, in the significant gray area band, delivered with the specific judicial language that produced it. That grounding is what makes the output resistant to hallucination: the score is tied to extant text, not to model probability. The paper presents the example as a demonstration of design logic, not as an evaluated result.

Validation strategy and falsifiability

How do you validate a system whose correct answer is that there is no correct answer? The paper builds ground truth from three sources: formal acknowledgments such as documented circuit splits and certiorari grants, expert ratings from judges and professors on a five-point scale, and outcome signals such as high reversal rates. Testbeds are chosen to exercise different types: employment discrimination for normative indeterminacy, arbitration for methodological, the Fourth Amendment for jurisdictional. Success requires a Spearman correlation above 0.7 with expert judgment and inter-rater reliability (Cohen's kappa) above 0.6. Failure is declared if correlation falls below 0.3 or if experts find more than 20 percent of the cited conflicts spurious. The framework must also beat three simpler baselines: a keyword heuristic, a citation-network measure and direct prompting of a language model. The framework awaits this empirical test; its contribution so far is theoretical and architectural.

Conclusion: from oracle to cartographer

The paper closes on a change of role. An AI that acts as an oracle hides the law's open texture behind a single answer. An AI that acts as a cartographer marks where the ground is contested, and in a profession built on resolving conflict, knowing what is not settled is itself a form of legal intelligence.

Where to go next

Frameworks in this piece

Terms in this piece

Revision history

23 Jul 2026Paper page created.

How to cite

Brodskiy, R. (2026, July 23). Detecting Genuine Doctrinal Ambiguity: A Multi-Layer Framework for Identifying Structural Indeterminacy in Judicial Reasoning [Working paper]. SSRN. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7039398

Related pieces

Follow the research

New papers, frameworks and essays. No marketing. Or use RSS.