Detecting Genuine Doctrinal Ambiguity
A Multi-Layer Framework for Identifying Structural Indeterminacy in Judicial Reasoning

Current AI systems suffer from 'false confidence,' presenting contested legal conclusions as settled. While computational law has modeled reasoning, a gap exists in detecting where reasoning breaks down due to structural indeterminacy. This article addresses this by distinguishing between 'shallow ambiguity' (resolvable by research) and 'genuine ambiguity' (structural features of law). We propose a five-type taxonomy (semantic, normative, methodological, jurisdictional, and analogical) accompanied by the Gray Area Detection Framework. This three-layer architecture utilizes extraction-time signals, query-time aggregation, and structural analysis to identify relational ambiguity across large corpora. By shifting focus from outcome prediction to contestation detection, the framework mitigates LLM-era risks like hallucination and decontextualization. We specify a validation methodology using expert annotation, formal acknowledgments, and domain-specific testbeds (Employment Discrimination, Arbitration, Fourth Amendment) to measure system utility. Ultimately, this approach moves legal AI from an 'oracle' model toward a systematic 'cartographer' of the law's gray areas, enhancing judicial transparency and reinforcing the rule of law through verifiable explanations of doctrinal conflict.
- Authors
- Ross Brodskiy
- Posted
- 23 July 2026 (SSRN)
- SSRN abstract
- 7039398
- Length
- 13 pages
- DOI
- Not yet assigned
- Keywords
- legal indeterminacy, computational jurisprudence, doctrinal conflict, gray area detection, legal AI safety, case-based reasoning
- Shallow ambiguity is a research problem that more searching resolves; genuine ambiguity is a structural feature of the law that no amount of research removes.
- An operational definition sidesteps the Hart and Dworkin debate: a question is genuinely ambiguous when the legal system itself treats it as contested, through splits, hedging or vigorous dissent.
- Genuine indeterminacy comes in five detectable types (semantic, normative, methodological, jurisdictional and analogical), each tied to a theoretical source and a distinct detection strategy.
- Ambiguity is relational: a circuit split is invisible in any single opinion. Detection therefore needs three layers (extraction-time signals, query-time aggregation, structural analysis), implemented as 18 tests feeding a composite Gray Area Score.
- Every score is grounded in citable judicial language, so an attorney can check the conflicts the system reports instead of trusting a model's confidence.
- The framework is built to be falsified: it must reach a Spearman correlation above 0.7 with expert judgment, beat keyword, citation-network and direct-prompting baselines, and is treated as failed if correlation falls below 0.3 or more than 20 percent of cited conflicts prove spurious.
- Introduction: The Problem of Uncharted Gray Areas
- Jurisprudential Foundations: Defining Genuine Ambiguity
- Why Existing Computational Approaches Fail to Detect Indeterminacy
- The Multi-Layer Detection Architecture
- Validation Strategy and Falsifiability
- Conclusion: Toward a Systematic Jurisprudence
The paper in brief
Legal AI has spent most of its effort on getting answers right. This paper argues that the more urgent missing capability is knowing when there is no single right answer to get. When a system predicts one outcome for a question that has divided the federal circuits for decades, it is not being accurate. It is displaying false confidence, and in doing so it quietly decides a contested question in a black box. Brodskiy proposes to shift the goal from prediction to contestation detection: finding the places where the legal materials themselves fail to yield a unique outcome.
Introduction: the problem of uncharted gray areas
The paper calls indeterminacy detection the "missing layer" of the legal AI safety stack. Treating unsettled law as settled is more than a technical error. It fails to respect the dialectical character of common-law reasoning and erodes the transparency on which judicial legitimacy depends. The paper makes four contributions: a distinction between shallow and genuine ambiguity, a five-type taxonomy of indeterminacy, a three-layer detection architecture, and a validation program designed so the framework can fail.
Jurisprudential foundations
The central conceptual move is the shallow and genuine distinction. Shallow ambiguity comes from incomplete information: a question of first impression in one jurisdiction that other jurisdictions have already resolved is only shallowly open, because the answer exists and has not yet been found. Genuine ambiguity comes from the structure of law itself: the materials support incompatible readings, or the system treats the question as contested, as in a persistent circuit split. More research cannot resolve it.
The paper then turns legal philosophy into checkable requirements with a five-type taxonomy.
| Type | Theoretical source | What it looks like in opinions |
|---|---|---|
| Semantic | Hart's open texture | Courts construe the same term in divergent ways |
| Normative | Dworkin's hard cases, Alexy's balancing | Competing principles resolved "on balance" or by unweighted multi-factor tests |
| Methodological | Llewellyn's dueling canons, Eskridge, Grove | Majority and dissent invoke opposing interpretive canons |
| Jurisdictional | Federal circuit splits | Courts expressly acknowledge that the circuits are divided |
| Analogical | Sunstein, Levi | Courts treat different precedents as controlling on similar facts |
A reader steeped in jurisprudence will ask whether the project assumes Dworkin is wrong. The paper's answer is an operational stance. Even if Hercules could find a right answer in every hard case, real legal systems work under constraints that make such answers practically inaccessible. Defining genuine ambiguity as whatever the legal system itself treats as contested keeps the criterion neutral between positivism and interpretivism while remaining computable. A longer companion treatment of the same framework adds that categories can overlap (normative and methodological indeterminacy can blur), and that a question triggering several detectors is signaling compound indeterminacy, not exposing a flaw in the taxonomy.
Why existing computational approaches fail
Each earlier tradition models law on the assumption that its inputs are settled. The symbolic tradition of Allen and McCarty formalizes rules precisely, but it requires someone to resolve ambiguity before the logic can be written. Case-based reasoning in HYPO and CATO shows how to argue within a doctrine; it cannot tell when the doctrine itself is fragmenting. Argumentation frameworks in the Prakken and Sartor line formalize how rules defeat one another but presuppose that priorities between them can be specified. Legal NLP can label rhetorical roles, while large language models add hallucination and decontextualization, which the paper reads as indeterminacy failures in disguise: a model resolving a genuine tension in its training data with one confident, false statement. The summary line is that existing systems model law, and none of them detects where law runs out.
The multi-layer detection architecture
Ambiguity is rarely a property of a single document. It lives in the space between cases, and a circuit split appears only when opinions are mapped against their peers. The architecture therefore runs in three layers.
- Layer A, extraction-time signals (8 tests), read individual opinions: explicit split acknowledgment, first-impression declarations, judicial hedging, dueling authorities, dissent vigor, multi-factor balancing, canon conflict and precedent instability.
- Layer B, aggregate tests (6 tests), look across the corpus: stance inversion, jurisdictional divergence, temporal drift, outcome variance, dissent rate and reversal rate.
- Layer C, structural tests (4 tests), examine the doctrine itself: fragmentation, test-standard mismatch, constitutional penumbra, and the composite Gray Area Score that normalizes and weights every contributing signal.
The ordering answers a practical constraint. Real systems ingest opinions one at a time, so single-case signals are captured at extraction and confirmed by aggregation at query time. The companion treatment sets out score bands running from clear law through significant gray area to profound indeterminacy, and explains why the design prefers missed signals to false positives at extraction.
The paper's worked illustration is Collins v. Virginia (2018), on whether police need a warrant to search a vehicle parked in a home's driveway. Layer A picks up the majority's express reservation of questions it did not decide and a vigorous dissent predicting that the rule will prove unworkable. Layer B looks for divergence among lower courts applying the decision, and Layer C for the curtilage concept splintering into sub-rules. The illustrative result is a Gray Area Score of about 0.65, in the significant gray area band, delivered with the specific judicial language that produced it. That grounding is what makes the output resistant to hallucination: the score is tied to extant text, not to model probability. The paper presents the example as a demonstration of design logic, not as an evaluated result.
Validation strategy and falsifiability
How do you validate a system whose correct answer is that there is no correct answer? The paper builds ground truth from three sources: formal acknowledgments such as documented circuit splits and certiorari grants, expert ratings from judges and professors on a five-point scale, and outcome signals such as high reversal rates. Testbeds are chosen to exercise different types: employment discrimination for normative indeterminacy, arbitration for methodological, the Fourth Amendment for jurisdictional. Success requires a Spearman correlation above 0.7 with expert judgment and inter-rater reliability (Cohen's kappa) above 0.6. Failure is declared if correlation falls below 0.3 or if experts find more than 20 percent of the cited conflicts spurious. The framework must also beat three simpler baselines: a keyword heuristic, a citation-network measure and direct prompting of a language model. The framework awaits this empirical test; its contribution so far is theoretical and architectural.
Conclusion: from oracle to cartographer
The paper closes on a change of role. An AI that acts as an oracle hides the law's open texture behind a single answer. An AI that acts as a cartographer marks where the ground is contested, and in a profession built on resolving conflict, knowing what is not settled is itself a form of legal intelligence.
Where to go next
- The Grayness Score: the scoring framework stated on its own, with worked application.
- Measuring Doctrinal Indeterminacy: a scholarly retelling of the gray-area program.
- When the Law Is Not Settled: the same idea for readers outside the profession.
- Ariadne's Thread: the later program that measures openness as dispersion across judicial postures.
Frameworks in this piece
Terms in this piece
Revision history
| 23 Jul 2026 | Paper page created. |
How to cite
Brodskiy, R. (2026, July 23). Detecting Genuine Doctrinal Ambiguity: A Multi-Layer Framework for Identifying Structural Indeterminacy in Judicial Reasoning [Working paper]. SSRN. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7039398
Related pieces
New papers, frameworks and essays. No marketing. Or use RSS.


