Z, Formally
Inputs, Rules, Inference Procedure, Outputs: What Ariadne's Metric Measures and What Its Theorems Prove

"Does this case have a right answer?" is usually treated as one question. It is at least four. What inputs does the decider receive? What rules map those inputs to a disposition? What procedure infers the disposition from the rules? What output is returned, and to whom?
The Hart-Dworkin debate is conducted almost entirely about the second item, and it has not converged. Brodskiy and Pokov's Ariadne's Thread is interesting because it does something different with each item. It holds the inputs fixed across deciders. It varies the rules in a declared, controlled way. It specifies the inference as a variance decomposition. And it restricts the output to a measurement. This note restates the method in that order, then states what the formal results do and do not establish.
1. Definitions
Six terms carry the argument. They should be read as defined, not as ordinary English.
- Posture. A filter and an ordering over a fixed, shared list of legal moves: which moves are admitted, how they rank in conflict, and at what level of generality the controlling principle is stated. The paper is explicit that posture is a scholarly construct, not a term of binding authority.
- Gate. A point in a civil action at which a trial judge can shape or end the case. The paper fixes nine: pleading screen (G1), threshold exits (G2), discovery scope (G3), discovery enforcement (G4), class certification (G5), expert admissibility (G6), summary judgment (G7), trial gate (G8), remedy and damages scope (G9).
- Score. An integer from 1 to 10 at one gate, in a fixed screening direction: 1 means the claim almost certainly passes, 10 means it is almost certainly screened out. The score is the atomic datum of the method.
- Replicate. The same posture, the same case and the same instruction, run again as an independent instantiation. Replicates do not add postures; they measure the instrument.
- Noise floor. The pooled within-posture variance of replicate scores: the spread the instrument produces when nothing about the posture has changed.
- Legal openness. The portion of the disposition that changes when the posture changes and the legal substrate does not.
The last definition is the one to hold onto. Openness is not difficulty, not ambiguity of wording, and not unpredictability of a named judge. It is sensitivity of the disposition to lawful posture, with the law held constant.
2. What is and is not being measured
The object of measurement is narrow. Z at a gate estimates how much of the variation in scores across a declared panel of postures is systematic, as opposed to run-to-run inconsistency.
Several things are deliberately outside the measurement:
- Correctness. No term in the metric designates a legally correct disposition.
- The winner. Z does not say who should prevail. A low-Z gate can be one the claimant clearly loses.
- A named judge. The paper disclaims predicting what a particular judge will do with certainty.
- Advocacy quality. How well a position is argued is reported separately and never folded into Z. A poorly argued open case is still open; a brilliantly argued closed case is still closed.
- Ideology. The judge's ideological prior is held on a separate shelf from the dispositional axes. The paper's phrase is that disposition is the gun and the prior is its aim.
3. Inputs
A run requires the following, fixed before any score is produced:
- A case packet. One operative complaint and one motion under test.
- A verified substrate. For each load-bearing proposition, grounding in primary authority, current good-law status, and the jurisdiction and stage to which it speaks. The substrate must be identical for every posture. The authors supply it through a proprietary system, Solon, described only by this contract; the theory is substrate-agnostic.
- A validity gate result. Propositions resting on overruled, fabricated or misread authority keep the dispute out of measurement. Z is not defined for a defective dispute. A valid but aggressive reading, a stretch, is admitted with a flag. The authors implement the gate through a second proprietary system, Basanos, again described only by role.
- A declared panel. P postures placed deliberately to span the posture space, not sampled at random. At the trial tier, each is a setting of four conjectured axes: termination mass, termination location, process volume and process steering, with the ideological prior, standard of review and case knowledge held apart.
- Design parameters. The replicate count R and the abstention level alpha.
Item 2 is not an implementation detail. If each posture reasons over its own partly mistaken picture of the law, the spread measures differential legal error. Posture must be the only thing that varies.
4. Rules
At each implicated gate, each posture does three things.
- It returns a score, with a one-line rationale, or a declared decline-to-reach where its posture would not get there. Forcing every posture to score every gate would manufacture dispersion at gates that would never be reached.
- It states the standard of review at that gate, which feeds the durability weighting.
- It flags any point where its rationale presents a discretionary setting as a command of the law. These flags form the anti-pretext ledger.
The score channel supplies the measurement. The rationale channel supplies the audit. They are kept separate.
5. Inference procedure
The model is the one-way random-effects components-of-variance model, the standard tool of generalizability theory. Fix a gate g.
Setup
P postures in the declared panel, p = 1..P
R same-profile replicates per posture, r = 1..R
s(p, r, g) integer score in {1, ..., 10}
Model
s(p, r, g) = mu(g) + alpha(p, g) + eps(p, r, g)
Var(alpha) = sigma2_posture(g) systematic between-posture variance (target)
Var(eps) = sigma2_noise(g) replicate noise (the floor)
Observed dispersions
sbar(p, g) = (1/R) * sum over r of s(p, r, g)
sbarbar(g) = (1/P) * sum over p of sbar(p, g)
V_between(g) = (1/(P - 1)) * sum over p of (sbar(p, g) - sbarbar(g))^2
V_within(g) = (1/(P * (R - 1))) * sum over p, r of (s(p, r, g) - sbar(p, g))^2
Estimator and metric
sigma2_posture_hat(g) = V_between(g) - V_within(g) / R
Z(g) = min(1, max(0, sigma2_posture_hat(g)) / 20.25)
where 20.25 = (10 - 1)^2 / 4
Abstention
F(g) = R * V_between(g) / V_within(g)
df = (P - 1, P * (R - 1))
report Z(g) only if F(g) exceeds the upper-alpha critical value; else abstain
Aggregation over the gates G(c) a case implicates
Z(c) = sum of w(g) * Z(g) / sum of w(g), w(g) nonnegative, not all zero
Z_durable(c) = the same average, weight restricted to deferential gates
Three features of the procedure deserve comment.
The correction term. Each posture mean averages only R noisy replicates, so the variance of posture means contains the posture variance plus the noise variance divided by R. Subtracting V_within over R removes exactly that residue. The correction shrinks as R grows, which is why the naive and corrected numbers nearly coincide with many replicates and diverge, with the corrected form mandatory, in any affordable panel.
The normalizer. By Popoviciu's inequality, a quantity confined to an interval of width 9 has variance at most 81/4, attained only by equal mass at the two endpoints. Using this theoretical maximum rather than an empirical one means Z needs no reference corpus. A single motion no model has seen can be scored.
The aggregate. The case-level value is a convex combination of gate values, with weights tied to the standard of review. De novo gates, paradigmatically the pleading screen and summary judgment, are path-determinative but outcome-provisional. Deferential gates, such as discovery, class certification, expert admissibility and remedy scope, are outcome-durable. Restricting weight to the deferential gates yields Z_durable, and the deferential gate with the highest screening propensity is the durable kill point.
6. Outputs
A conforming output contains the per-gate Z or an abstention, an interval, the noise floor, the postures driving dispersion, flags on stretched propositions, a durability-weighted gate map, and the rationale and anti-pretext ledger. The later versions of the paper make the prohibitions explicit: no binding disposition, no disguised recommendation to a judge, and no bare score detached from its panel and variance ledger.
7. Why abstention is a feature
A system that always emits a number is a system that sometimes emits noise with two decimal places. The abstention rule makes the output three-valued: open, closed, or unmeasured.
The paper's illustrative numbers, labeled as illustrative, show the rule working. Eight postures, four replicates each, confront a complaint whose answer is nearly overdetermined: V_between is 0.30 and V_within is 0.50. The corrected variance is 0.175, so Z is about 0.009, and F is 2.4 on (7, 24) degrees of freedom, which does not clear a conventional threshold. The instrument abstains. Perturb the facts by one value-laden clause and let the panel split toward the endpoints: V_between rises to 9.0 with the same floor, Z is about 0.44, F is 72, and the reading is confident. Thin signal yields silence; a real split yields a number.
Abstention also disciplines the designer. The later versions report a logged smoke test on an engineered New York motion to dismiss: five posture scores of 2, 2, 4, 7 and 8, so V_between is 7.8, against a floor of 0.167 estimated from two replicated control profiles. Because the panel postures ran once each, the V_within over R correction could not be applied as written; the authors subtracted the whole floor and reported Z of about 0.38. They also reported that the formal abstention test could not be run, since an unreplicated panel has zero within-posture degrees of freedom, and that the sampling interval ran from roughly 0.14 to 1. That is the correct way to report a deviation from one's own protocol.
Two further cautions from the same versions apply. At small designs the test is weak: with five postures and three replicates, power to detect a posture component equal to the noise component is only about one half, so sensitivity must be bought with replicates. And reporting only on rejection inflates values near the threshold, which is why Z travels with an interval and a published abstention level.
8. What the theorems prove
The formal results, stated as theorems, concern the instrument:
- Unbiasedness. Under the model, V_between minus V_within over R has expectation equal to the systematic posture variance.
- Boundedness. The population value of Z lies in [0, 1], and the normalizer is tight.
- Predictive irrelevance. Adding the deciding posture to a predictor that already has the case and the law reduces expected squared error by exactly the posture variance. Population Z is zero if and only if posture is predictively irrelevant at the gate, which is the operational meaning of an easy gate.
- Aggregation. The case-level Z is bounded and nondecreasing in each gate value, so it cannot manufacture openness no gate exhibits.
- Level control. Abstaining unless F exceeds the upper-alpha critical value bounds the false-openness rate at alpha; a permutation test gives the same guarantee without the normal approximation.
A further structural result needs no distributional assumption: Z is invariant to which disposition is labeled correct and to a common shift in all scores. Objections to a language model's competence as a judge operate on a channel the metric does not read. The one objection that reaches the metric is that the postures do not really differ.
9. What the theorems do not prove
A theorem about a ruler is silent about the object it is laid against. None of the following follows from the results above:
- that the four axes are real and separable (C-Basis);
- that they predict dispositions beyond an ideology proxy (C-Increment);
- that Z tracks the hardness experts and dockets recognize (C-Zone);
- that the systematic component dominates noise often enough to be useful (C-Signal);
- that posture matters most at deferential gates (C-Deference);
- that the substrate is faithful enough to remove posture-correlated legal error (named C-Ground in the later versions).
Three qualifications sit inside the formal layer itself. Flooring at zero reintroduces a small positive bias near zero, so unbiasedness belongs to the unfloored correction. The model assumes homoscedastic replicate noise, which must be checked for a language-model instrument. And the numerator is panel-relative: a panel of near-clones reads low on any motion. A Z reported without its versioned panel is not a measurement.
The sharpest open risk is separability. Recent experiments report that frontier models behave as one formalist persona resistant to steering. Grounding the substrate does not solve this and can hide it: identical personas over identical correct law converge, and the resulting low Z is indistinguishable from determinacy unless a clone discriminator, dispersion on motions experts rate hard, is part of the validation.
The specification is complete enough to implement and to refute. That is the right state for a measurement proposal to be in. Whether the ruler has been laid against the right object is an empirical question, and the paper names the experiment that answers it.
Frameworks in this piece
Terms in this piece
Revision history
| 18 Jul 2026 | First published in the Institute library. |
How to cite
Mercer, A. (2026, July 18). Z, Formally: Inputs, Rules, Inference Procedure, Outputs: What Ariadne's Metric Measures and What Its Theorems Prove. Computational Law Institute. https://institute.legawrite.ai/articles/ariadne-z-formally
Related pieces
New papers, frameworks and essays. No marketing. Or use RSS.


