New on SSRN: Ariadne's Thread, a measurement-theoretic method for legal openness. Read the paper
FrameworkDeterminacy and Gray AreasdraftVersion 0.1 · 23 Jul 2026

Grayness Score and Gray Area Radar

The Grayness Score is a composite, evidence-linked indicator of how far the legal system itself treats a question as contested, and the Gray Area Radar is the view that shows a lawyer which signals fired and where they came from.

Origin: Detecting Genuine Doctrinal Ambiguity · Detection layers and composite score specified in Detecting Genuine Doctrinal Ambiguity (SSRN, posted 23 July 2026), where the score is called the Gray Area Score.

1Clear Law2345Profound Indeterminacy
1
Clear Law Draft band 0.0 to 0.2; treated as settled
2
Minor Uncertainty Draft band 0.2 to 0.4; weak or isolated signals
3
Moderate Ambiguity Draft band 0.4 to 0.6; mixed signals across layers
4
Significant Gray Area Draft band 0.6 to 0.8; contest visible and grounded
5
Profound Indeterminacy Draft band 0.8 to 1.0; materials do not fix outcome
The five tiers proposed in the working draft, from settled law to profound indeterminacy; the bands are provisional until the weights are calibrated.

The framework

Legal AI tends to fail on unsettled questions in a particular way. It finds real authority for one side, stops looking, and answers with confidence. The citations check out. The trouble is that the question divides the courts, and nothing in the answer says so. The Grayness Score and the Gray Area Radar are the Institute's names for the instruments meant to catch this. They are built from the detection framework in Detecting Genuine Doctrinal Ambiguity and the working drafts behind it.

Two kinds of uncertainty

  • Shallow ambiguity is a research problem. Controlling authority exists but has not been found, or an apparent conflict dissolves on a closer read. More research fixes it.
  • Genuine ambiguity is structural. Courts reach conflicting conclusions on the same question, authoritative sources support incompatible readings, or principles of similar weight pull in opposite directions. No amount of research fixes it.

The test is operational: a question is genuinely ambiguous when the legal system itself treats it as contested, through acknowledged splits, hedging opinions, vigorous dissents, or outcomes that vary unpredictably. This sidesteps the Hart and Dworkin debate. Even if an ideal judge could find a right answer, a question that working courts treat as contested is contested for the lawyer advising a client.

Genuine ambiguity comes in five types, each with its own detection strategy: semantic (Hart's open texture: courts construe the same term differently), normative (principles of similar weight balanced differently), methodological (majority and dissent reaching opposite results through competing interpretive canons), jurisdictional (courts of equal authority split), and analogical (courts treating different precedents as controlling for similar facts). The boundaries blur, normative with methodological and semantic with analogical. A question that triggers several types signals compound indeterminacy rather than a flaw in the taxonomy.

Three detection layers

Ambiguity is relational. A circuit split is invisible inside one opinion and appears only when opinions are compared. Real systems ingest cases one at a time, so detection runs in phases: extract hints from each opinion, then confirm them across the corpus.

LayerWhen it runsTests
A. Extraction-timeAs each opinion is ingestedA1 explicit split acknowledgment; A2 first-impression declaration; A3 judicial hedging; A4 dueling authorities; A5 dissent vigor; A6 multi-factor balancing; A7 canon conflict; A8 precedent instability
B. AggregateAt query time, across the corpusB1 stance inversion; B2 jurisdictional divergence; B3 temporal drift; B4 outcome variance; B5 dissent rate; B6 reversal rate
C. StructuralOver the doctrine as a wholeC1 doctrinal fragmentation; C2 test-standard mismatch; C3 constitutional penumbra; C4 composite score

Extraction favors precision over recall. A missed signal is preferred to a false one that would propagate into every aggregate query built on it.

The Grayness Score

The Grayness Score is the Institute's name for test C4, which the paper calls the Gray Area Score: a weighted combination of normalized signals. Binary signals enter directly, counts logarithmically, and ratios linearly against configurable thresholds. The computation is deterministic and reproducible. The weights, however, are configurable parameters to be calibrated through validation. No source fixes them, and this page does not either. That is why the framework carries draft status.

The working draft proposes five tiers in equal bands from 0 to 1 (see the diagram). An earlier draft proposed a different five-point confidence scale, and the two have not been reconciled.

The Gray Area Radar

A number alone would reproduce the problem it is meant to solve: one more confident-looking output. The Radar is the view that makes the score inspectable. For a given question it shows which tests fired, grouped by layer and by ambiguity type; the passages and cases behind each signal; and whether the uncertainty is shallow ("not found yet") or genuine ("courts disagree"). The source framework requires exactly this grounding, so that an attorney can open the cited cases and confirm that the conflict is real. The sources leave practitioner interfaces as an open design problem. The Radar is the Institute's specification of that interface, not a shipped product.

How to apply it

The sources trace one question through the layers, and they present it as a design illustration, not an evaluated result. Does the Fourth Amendment require a warrant before police search a vehicle parked in a home's driveway?

The question sits where two doctrines meet: the automobile exception, and the curtilage doctrine, which extends home-like protection to the area around a residence. In Collins v. Virginia (2018) the Supreme Court held that the automobile exception does not permit warrantless entry into curtilage to search a vehicle. Justice Alito dissented, arguing that the result was inconsistent with prior cases and that the rule would prove unworkable.

On the Radar, the expected signals are:

  • Layer A. A1: Collins resolved a split, but lower courts remain unsure how far it reaches. A3: the sources read the majority as reserving questions about other settings, such as apartment parking areas. A5: dissent vigor is high. A8: the majority distinguishes several automobile-exception precedents without a clear limiting principle.
  • Layer B. B2: lower courts diverge in application, some extending Collins broadly and others confining it to enclosed spaces. B5: an above-average dissent rate on curtilage boundary questions.
  • Layer C. C1: "curtilage" fragments into sub-rules for different architectural settings.

The dissertation draft expects a score between 0.5 and 0.7, straddling Moderate Ambiguity and Significant Gray Area; the SSRN version uses 0.65 as its illustration. Either way, the Radar's job is the same: cite Collins, quote the reserved questions and the dissent, and list the divergent lower-court decisions so the lawyer can check each one.

What the lawyer does next depends on the reading. Shallow uncertainty calls for more research. Here the uncertainty is genuine, and the gray-area essays prescribe a different response: identify the conflict, map which courts say what, flag it explicitly, present the competing positions with an honest view of which is stronger, and counsel the client that the question is contested. In a brief, that means acknowledging the adverse line and arguing why the client's side is better, not writing as if the question were settled. The persona essays When the Law Is Not Settled and Measuring Doctrinal Indeterminacy develop this for practitioners and scholars.

Known limitations and critiques

No validation yet. The framework is theoretical and architectural. Its proposed protocol sets its own bar: a Spearman correlation around 0.7 or higher with expert ratings in at least two domains, expert agreement above a Cohen's kappa of 0.6, outperformance of keyword, citation-network, and direct-prompting baselines, ablations showing that each layer contributes, and more than 80 percent of cited conflicts confirmed as real on expert review. It counts itself failed if the correlation falls below 0.3 or more than 20 percent of cited conflicts prove spurious. The sources report no results against any of these.

Uncalibrated numbers. Without weights, the tier bands are placeholders, and the two drafts' scales disagree. A score of 0.65 today is an illustration, not a measurement.

The ground truth is itself contested. Validating a detector of disagreement requires experts to agree about disagreement. Formal acknowledgments such as certiorari grants help, but they lag the law they describe.

Courts do not always announce doubt. Layer A depends on judges signaling uncertainty in their language. A court that papers over a real conflict leaves no hedging to extract, and a judge who hedges by habit produces signals that mean little. The aggregate layer is meant to catch both cases, but it inherits whatever extraction missed.

Some tests carry a jurisprudential position. Treating rights derived from structural inference as a marker of indeterminacy (C3) is a contestable view of constitutional law, not a neutral measurement.

A snapshot, not a forecast. The score describes the current state of doctrine. It cannot say whether a split will persist or how it will resolve. It is also built for English-language common law; civil-law systems would need adaptation.

Knowledge acquisition is the bottleneck. The aggregate tests need holdings extracted, stance-classified, and linked at corpus scale. Outside narrow domains that remains expensive.

Lexicon terms

Related frameworks

Pieces that use this framework

Changelog

v0.1 · 23 Jul 2026Three detection layers, eighteen tests and a composite score specified; weights and thresholds await calibration.

How to cite

Computational Law Institute (2026, July 23). Grayness Score and Gray Area Radar (Version 0.1). https://institute.legawrite.ai/frameworks/grayness-score

Cite version 0.1; the changelog above records what changed.