Ariadne's Thread
Legal Determinacy as Measured Dispersion across Posture-Instantiated Adjudicators

Ronald Dworkin's Hercules is the ideal judge who reads the whole legal tradition as one coherent scheme of principle and rules the single right answer. No one has built him, and as an ideal he carries two liabilities: he decides, and he presupposes that a determinate answer exists. This paper proposes a corrected ideal, Ariadne, who guides rather than rules: the dispute is the labyrinth, the judicial postures available on the question are its corridors, and the metric Z is the thread. The contribution is a single move made rigorous. The question whether a contested case has a single right answer is converted from a philosophical premise into an empirical quantity. A panel of posture-instantiated judge-agents scores a procedural motion on a fixed grid of gates, figure-skating style, one integer per gate with a one-line rationale; Z is the dispersion of their predicted dispositions between postures, net of a same-profile replicate noise floor, normalized to a bounded scale. Low Z means the merits overdetermine the outcome; high Z means the judge's posture is the tie-breaker; noise of the same order as signal means the instrument abstains. Z is given a formal foundation in the components-of-variance model, and the theorems proved about it (that the noise-corrected estimator is unbiased, that Z is bounded by a tight normalizer, that a gate has zero dispersion exactly when the deciding judge's posture is predictively irrelevant there, that the case-level aggregate is a bounded, monotone, durability-weighted combination of gate values, and that abstention controls the rate of falsely declaring a noise gate open) are properties of the ruler, not facts about what it measures. The measurement reports openness only when every posture reasons over the same verified, good-law substrate. The facts about judges and language models, including the strongest threat that frontier models behave as one rigid formalist persona, are stated as falsifiable conjectures and deferred to a calibration experiment on which the program is conditional. The claim is held to procedural dispositions in large-dollar civil litigation, deployed as a measurement-and-audit instrument under human control.
- Authors
- Ross Brodskiy and Nathan Pokov
- Posted
- 18 July 2026 (SSRN)
- SSRN abstract
- 7001179
- Length
- 33 pages
- DOI
- Not yet assigned
- Keywords
- legal determinacy, indeterminacy thesis, Dworkin, Hart, judicial behavior, components of variance, generalizability theory, computational law, artificial intelligence and law, measurement
- Legal determinacy can be treated as a measurable property of a case: the systematic dispersion of qualified, differently disposed adjudicators confronting the same motion, rather than a premise the Hart-Dworkin debate must settle first.
- Z is the between-posture variance of gate scores minus the same-profile replicate noise floor (V_between minus V_within over R), divided by 20.25, the largest variance a quantity on the 1 to 10 scale can have; it sits on a fixed 0 to 1 scale and can be computed for a single new motion, and the instrument abstains when the spread does not clear the floor.
- Z is zero at a gate exactly when the deciding judge's posture adds nothing to prediction of the disposition, which is the operational meaning of an easy gate for positivists and Dworkinians alike: the number is the same and only the gloss differs.
- Durability matters: rulings at deferential gates (discovery, class certification, expert admissibility, remedy scope) tend to survive appeal, so the method reports a durability-weighted Z and a durable kill point alongside the path-level aggregate.
- Dispersion measures openness only when every posture reasons over one verified, good-law substrate; grounding solves substrate reliability but not posture separability, and identical personas over identical law would produce a clean low Z that looks like determinacy.
- The theorems are properties of the ruler; whether postures are separable from ideology, whether language models can instantiate them, and whether Z tracks real case-hardness are falsifiable conjectures deferred to a calibration experiment on which the whole program is conditional.
- I. Introduction: From Hercules to Ariadne
- II. The Problem the Method Solves
- III. Two Posture Families, One Architecture
- IV. The Mechanism: The Swarm and the Figure-Skating Scores
- V. The Formal Model and Its Theorems
- VI. The Durability Law
- VII. The Anti-Pretext Output
- VIII. A Worked Example
- IX. Novelty and Prior Art
- X. The Empirical Conjectures and Their Refutation Tests
- XI. The Validation Gate, as Future Work
- XII. Limitations and Legitimacy
- XIII. Conclusion
- Appendix A. Proofs and Derivations
- Appendix B. Provenance and Authorities Note
The paper in brief
From Hercules to Ariadne
Ronald Dworkin's Hercules is legal theory's most demanding ideal: a judge of superhuman skill who reads the whole legal record as one coherent scheme of political morality and rules the answer that reading yields. No one has built him, and the artificial intelligence and law literature has not tried to; its systems model fragments of legal reasoning, from case-based argument to outcome prediction.
Brodskiy and Pokov argue that Hercules is the wrong ideal for an instrument, for two reasons. He decides, so an instrument built in his image inherits every legitimacy objection to automated adjudication. And he presupposes the thing in dispute: his method makes sense only if a single best interpretation exists, which is exactly what the half-century argument between Hart and Dworkin has not resolved. The authors replace him with Ariadne, who in the myth neither fights the Minotaur nor walks the maze for Theseus but hands him the thread. The dispute is the labyrinth, the available judicial postures are its corridors, and the metric Z is the thread. Ariadne issues a measurement, not a verdict.
The thesis is that legal determinacy is a measurable property of a case. Instead of assuming an answer to the Hart-Dworkin question, the method measures how far qualified but differently disposed adjudicators diverge on the same motion. If they converge, the merits decide and the identity of the judge does not predict the outcome. If they diverge, the case is open and posture breaks the tie. If the divergence is no larger than the instrument's own unreliability, the instrument abstains. Run over a representative population of disputes, the distribution of Z would itself report which jurisprudential picture the data support. A positivist and a Dworkinian read the same number; only the gloss differs.
Posture as a construct
A posture, in the paper's usage, is a scholarly construct, not a term of binding authority. It is a filter and an ordering over one shared list of moves: text, precedent, purpose, structure, the canons, consequences and the rest. Named judicial philosophies differ not in the list but in which moves they admit, how they rank them in conflict, and at what level of generality they state the controlling principle.
There are two families. Meaning-fixing postures operate at the appellate tier, where the contested move is what a text means. Dispositional postures operate at the trial level, where the contested move is how a procedural motion is disposed of. The paper develops the dispositional family, whose rulings are observable, frequent and reviewable, and conjectures that a trial judge's dispositional posture is captured to first order by four axes: termination mass (how much the judge screens), termination location (where in the sequence the judge screens), process volume (how much discovery the judge authorizes) and process steering (how the judge steers contested process and admissibility questions). Three layers are held apart so they do not contaminate the measurement: the ideological prior, the standard of review, and case knowledge. The paper's image for the first separation is that disposition is the gun and the prior is its aim.
The gate grid and figure-skating scores
The shared move-list at the trial tier is a fixed grid of nine procedural gates: the pleading screen, threshold exits, discovery scope, discovery enforcement, class certification, expert admissibility, summary judgment, the trial gate, and remedy and damages scope.
Every agent in the panel receives the same complaint and the same motion. At each implicated gate it returns a single integer from 1 to 10, figure-skating style, where 1 means the claim almost certainly passes the gate and 10 means it is almost certainly screened out, with a one-line rationale in that posture's voice. It states the standard of review at the gate, and it flags any point where its own rationale presents a discretionary setting in the language of legal compulsion. The flags are collected into an anti-pretext ledger, which does not say a ruling is wrong. It says the ruling is a choice, and it locates the choice.
Z and its noise floor
The procedure is a noise audit in the sense Kahneman, Sibony and Sunstein describe: present many deciders with the same case and measure the spread of their judgments. The spread has two sources. Part of it tracks real differences in posture, and so tracks how open the question is. Part of it is the instrument's own inconsistency: the same profile, asked the same question again, will not always answer identically.
The paper models each score as a gate mean plus a posture effect plus replicate noise, the one-way components-of-variance model of generalizability theory. Each posture is run R times. The variance of the posture means overstates the systematic component by the noise variance divided by R, so the estimator subtracts exactly that: V_between minus V_within over R. Z at a gate is this corrected quantity, floored at zero and divided by 20.25, the largest variance any quantity confined to the 1 to 10 scale can have. Because the normalizer is theoretical rather than drawn from a corpus, Z can be computed for a single novel motion.
A case implicates several gates, so the case-level Z is a weighted average of gate values, with weights tied to the standard of review. Rulings at de novo gates, paradigmatically the pleading screen and summary judgment, are path-determinative but outcome-provisional. Rulings at deferential gates, such as discovery, class certification, expert admissibility and remedy scope, tend to stick. The method therefore reports the path-level aggregate, a durability-weighted Z, and the durable kill point: the deferential gate where a differently disposed judge most durably ends the case.
The theorems are properties of the ruler
The formal core proves that Z behaves like a measuring instrument. The noise-corrected estimator is unbiased for the systematic between-posture variance. Z is bounded in the unit interval by a tight normalizer. Z is zero at a gate if and only if the deciding judge's posture is predictively irrelevant there, which is the operational content of an easy gate on either jurisprudential reading. The case-level aggregate is bounded and monotone, so it cannot manufacture openness that no gate exhibits. And reporting Z only when an F test, or distribution-free a permutation test, rejects the null of pure noise controls the long-run rate of declaring a noise gate open. A further structural result shows that Z reads only the spread of posture means against the noise floor, never the correctness of any ruling, so objections to a language model as a judge do not bear on the metric itself.
The limit is explicit: the paper proves properties of the ruler, not facts about what the ruler measures.
Conjectures and refutation tests
The claims about the world are stated as named conjectures, each with the result that would refute it. C-Basis holds that the four axes are separable; strong cross-loading in a factor analysis would collapse the basis. C-Increment holds that the axes add predictive power beyond an ideology proxy; the paper calls it the bet most likely to be lost. C-Zone holds that Z tracks the case-hardness experts and dockets recognize. C-Signal holds that the systematic component dominates noise often enough for the instrument to be useful. C-Deference holds that posture matters most at deferential gates; equal reversal rates across the two clusters would refute it.
The strongest objection is stated candidly. Recent work reports that frontier models behave as a single rigid formalist persona that resists steering. If every agent is the same formalist under different labels, Z is an artifact and the program fails.
The verified-substrate precondition
Dispersion measures openness only if every posture confronts the same law. Agents that each hold a partly mistaken picture of the governing authority will diverge because they are arguing about different law. The method therefore requires one verified, good-law substrate shared by every posture, so that posture is the only thing that varies, and a validity gate that keeps disputes built on dead, fabricated or misused authority out of the measurement. The authors supply both through proprietary systems, Solon and Basanos, which the paper describes only by what they deliver. It is equally candid that grounding solves substrate reliability but not posture separability: identical personas over identical correct law converge harder, and their low Z would look like determinacy when it is in fact persona collapse.
Scope, and the calibration experiment
The claim is held to a narrow universe: procedural dispositions in large-dollar civil litigation at the trial level, deployed as a measurement-and-audit instrument under explicit human control. The paper does not claim to resolve whether right answers exist, to have built a general adjudicator, or that the instrument should ever decide anything.
The program is conditional on one experiment. Stage 1 runs the swarm over a held-out corpus of decided large-dollar civil motions spanning the gate grid and asks whether between-posture dispersion clears the noise floor by a pre-registered margin and correlates with observed inter-judge variation and with expert ratings of hardness. If the postures collapse, the program pivots, most plausibly toward activation-level steering. Stage 2 tests C-Increment and C-Deference against real dockets. Stage 3 ships Z only with its interval, its noise floor and the anti-pretext ledger, never as a bare score. Later versions of the paper add a logged smoke test on an engineered New York Commercial Division motion, in which five postures split three to two with Z of about 0.38 at the pleading gate, and label it for what it is: a demonstration of mechanics, not validation.
Variations and reviews
- Z, Formally: Ada Mercer restates the method as inputs, rules, inference procedure and outputs, with the estimator written out.
- The Thread in the Courtroom: Elena Voss tells the idea through an invented night before a motion to dismiss is heard.
- Before You File: Tom Brennan on what a low or high Z would change about budget, settlement and motion strategy, and what to ask before trusting one.
- From Hercules to Ariadne: Julian Fenwick places the paper in the intellectual history of the Hart-Dworkin debate.
- Priya Raman reviews Ariadne's Thread: the Scholar's review.
- Nora Kestrel reviews Ariadne's Thread: the Institutional Dissenter's review.
- Marcus Hale reviews Ariadne's Thread: the Common-Law Pragmatist's review.
- Eleanor Voss reviews Ariadne's Thread: the Doctrinal Architect's review.
- What Remains for the Judge When the Machine Has Already Verified the Law?: Ross Brodskiy's Russian-language Medium essay (20 June 2026) on the A-axes and the durable kill point.
- From the Supreme Court Machine to Ariadne's Thread: Ross Brodskiy's Russian-language Medium essay (21 June 2026), written for the continental lawyer.
Variations and reviews
Frameworks in this piece
Terms in this piece
Revision history
| 18 Jul 2026 | Paper page created. |
How to cite
Brodskiy, R., & Pokov, N. (2026, July 18). Ariadne's Thread: Legal Determinacy as Measured Dispersion across Posture-Instantiated Adjudicators [Working paper]. SSRN. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7001179
Related pieces
New papers, frameworks and essays. No marketing. Or use RSS.


