New on SSRN: Ariadne's Thread, a measurement-theoretic method for legal openness. Read the paper

What a Filing-Grade Benchmark Must Measure

Dispositive recall under doctrinal and procedural constraint, and why fabrication rate is no longer the headline

A sculptural tower of offset disks in clear glass, white marble and gold leaf rises from a round marble base around a softly glowing central column, set against a pale gray backdrop.
Plate 26 · Stacked LedgersPlates

Evaluations of legal research systems have, until recently, asked two questions. Did the system find a relevant case? Did it invent a citation? Both were reasonable questions when systems routinely fabricated authority. Both are now nearly exhausted. Grounded retrieval returns real cases; citation verification returns real citations. A benchmark that stops there can no longer distinguish among the systems worth distinguishing.

This guide sets out what the Institute takes a filing-grade benchmark to require. Its central claim is that the question which decides whether a brief can be filed is a third one, distinct from both integrity and recall:

Did the system surface the proposition that does the legal work, with the correct stance, at the correct procedural stage, with confirmed good-law status, in a form that is currently citable in this forum?

The guide proceeds in six steps: why fabrication must lose its place as the headline; the terms of art the standard depends on; six design principles; the dimensions to be scored; three archetypal test items; and the conditions under which results may be published. It closes with objections and a checklist.

1. Why fabrication is no longer the headline

A strictly grounded system, one whose citations are drawn from a real corpus and checked against primary sources, does not fabricate. Neither does a system that removes the language model from the citation path altogether. On the hard tier of any serious test, the only systems that fabricate are those that let the model emit citations. Fabrication therefore separates the ungrounded from the grounded. It does not separate grounded systems from one another.

The distinction that matters instead can be stated in one sentence:

Verified means the citation exists and is accurate. It does not mean the cited proposition is still controlling, on-stance, decided at the right stage, or currently citable.

A grounded system can return a real, correctly formatted opinion that is the wrong authority for the task, or miss the dispositive one entirely. Two failure types follow, and they deserve names:

  • True-but-useless retrieval: a real, accurately cited authority that cannot do the work asked of it (wrong stage, wrong forum, superseded, limited to its facts).
  • True-but-adverse retrieval: a real, accurately cited authority that in fact cuts against the party the lawyer represents.

Neither is a hallucination. Both pass every existing integrity check, which is precisely what makes them more dangerous than fabrication. A reversed intermediate-appellate opinion, an on-point holding from the wrong department, or a summary-judgment ruling offered on a motion to dismiss is not a weaker answer. It is an unusable answer that looks identical to a winning one.

The standard does not discard fabrication. It reports fabrication first, then holds it constant, so that comparison turns on what actually differs. This mirrors the Institute's broader position, developed in Good Law for What? and the Proposition-Usability Model, that validity is necessary for usability but far from sufficient.

2. Terms of art

Filing-grade. An authority is filing-grade for a proposition when a careful partner, checking before filing, would accept it: it states the proposition, it is binding or appropriately persuasive in the forum, it is good law, it has not been superseded or limited, it favors the client, and it was decided at a stage that makes it relevant to the present motion.

Dispositive recall under doctrinal and procedural constraint. The proportion of test questions on which a system returns the proposition that does the legal work while satisfying every stated constraint. This is the construct the benchmark exists to measure.

Dispositive Recall @ filing. The composite metric: a question is scored as passed only if a single answer is the right proposition, with the right stance, at the right stage, still good law, and currently citable. The composite is conjunctive at the level of the individual question.

Propositional answer. A gold answer specified as a rule, not a topic: the rule that says Y, established by court Z, still good law, favoring party P, citable in forum F. Topical proximity earns no credit.

3. Six design principles

  1. Select for structural failure, not topical difficulty. Test items should be chosen by task type, where the failure is architectural and cannot be repaired by a better prompt, a larger context window, or more retrieved passages. Obscure topics measure coverage, not design. If a dimension's gap could be closed by tuning, it does not belong in the hard tier.

  2. Score propositions, not topics. A system that retrieves "cases about X" when the question asks for "the rule that says Y" scores zero on that item, even if the right case appears somewhere in its results.

  3. Hold integrity constant so that recall is visible. Report fabrication, then set it aside. The comparison should turn on recall under constraint.

  4. Tie every dimension to a checkable property of the authority. Stance, subsequent history, stage, forum, and scope are properties that can be verified independently of any system. A system with no representation of, say, stance or subsequent history cannot filter on it by construction, and the benchmark should make that absence visible rather than incidental.

  5. Compose constraints conjunctively. Any system may clear one filter by luck. The bar a partner applies is that one answer clears all of them together.

  6. Use a real corpus with real controlling-law structure. The draft this guide adapts builds its test on New York case law, because that forum supplies genuinely hard controlling-law structure: splits among the Appellate Division's departments, subsequent history in the Court of Appeals, and the difference between holdings on a CPLR 3211 motion to dismiss and holdings on a CPLR 3212 motion for summary judgment. The framework is meant to generalize, and the draft includes a California variant turning on review and depublication and a federal circuit-split variant as illustrations.

4. What to measure

The dimensions fall into three clusters and a composite. Each row is a question a careful lawyer asks before filing.

ClusterDimensionThe question it asks
BaselineCitation integrityDoes the citation exist and say what is claimed?
BaselineTopical recallDid the system find any real, on-point authority?
ValidityGood lawIs it still controlling, not overruled or narrowed?
ValidityBinding versus persuasiveIs it binding in this forum, not merely on point?
ValidityHierarchyIs it the highest controlling court on the point?
ValiditySplit detectionDoes the system flag a department or circuit split and identify the controlling line?
ValiditySubsequent historyHas it been reversed, modified, vacated, or taken up on review?
ValidityStatutory supersessionHas a statute abrogated the common-law rule?
ValidityControlling statusIs it currently citable for the cited proposition?
ValidityMost recent controllingHas a newer case superseded the old landmark?
ValidityHolding scopeIs it good law but limited to facts that do not control here?
FitnessStanceDoes it favor my party?
FitnessHolding, not dictaIs the proposition the holding?
FitnessOriginDid this court establish the rule, or apply it?
FitnessStageIs it a pleading-stage holding when the motion is at the pleading stage?
FitnessMateriality standardDoes it state the controlling threshold?
FitnessCounter-authorityDid the system surface the authority that hurts?
FitnessDistinguishing patternIs the adverse authority distinguishable, and how?
FitnessOperative languageDid it return the operative rule, not the surrounding paragraph?
CompositeDispositive Recall @ filingDoes one answer satisfy proposition, stance, stage, good law, and citability together?

Three observations about this instrument. First, the baseline cluster is not decoration. Publishing topical recall, including where competing systems do well, is what makes any gap in the other clusters credible: it rules out the cheap explanation that one system simply has a larger corpus. Second, the validity cluster is where grounded systems are most exposed, because the facts that determine validity (subsequent history, review, supersession) are frequently absent from the text of the retrieved opinion itself. A system that reasons well over retrieved text cannot reason over facts the text does not contain. Third, the fitness cluster connects the benchmark to procedural posture, the failure examined in Right Law, Wrong Stage.

5. Three archetypal test items

The following archetypes illustrate what a hard-tier item looks like. Each is described by what a passing answer must do, not by any system's performance.

The department-split item. The question is posed in a Second Department forum. On the point in question, the First Department has adopted a rule the Second Department has not. The First Department's line is better known and more frequently cited, so a topical retriever will tend to surface it. Under New York stare decisis, as the draft states it, a trial court follows its own department where that department has spoken. A passing answer returns the controlling Second Department line, flags the split, and presents the First Department position as counter-authority to be briefed around.

The controlling-status item. The question asks for the strongest authority that a claim survives a motion to dismiss. A squarely on-point intermediate-appellate opinion exists. In the New York variant, it has since been reversed by the Court of Appeals; in the California variant, the state Supreme Court has granted review, which bears on its citability under rule 8.1115 of the California Rules of Court. The case is real, the citation is accurate, and the authority was once good. A passing answer catches the subsequent history, declines to offer the opinion, and returns the next currently citable authority.

The stage item. The question asks for the best authority that a fraud claim fails for want of particularity, on a CPLR 3211 motion. A real, correctly cited fraud decision exists holding the evidence insufficient, but it was decided at summary judgment on an evidentiary record. Offered on a motion to dismiss, where the question is particularity on the face of the complaint, it is useless, and citing it signals that the drafter does not know the standard. A passing answer confines itself to pleading-stage holdings.

In each case the failing answer is not a hallucination. It is a real case that fails on a structural axis.

6. Scoring the composite

The composite deserves care, because "multiplicative" can mislead. The standard does not multiply aggregate per-dimension rates; constraints in legal research are correlated, and a product of marginal rates would misstate the joint rate. It scores each question conjunctively and reports the share of questions on which a single answer clears every constraint. The drafting intuition is nonetheless right: when several independent-seeming constraints must all hold, the joint pass rate falls well below any single dimension's rate.

The draft offers a hypothesis about where that fall will be steepest. An agentic system built on a frontier model should outperform pure similarity search on stance and stage, because it can reason about them. It should still meet a ceiling on the validity cluster, because validity and controlling-law facts are often not present in retrieved text, and because a model that remains in the citation path retains some fabrication risk. That hypothesis is testable, and the benchmark exists to test it.

Two further scoring rules should be disclosed rather than assumed. The first is partial credit. The draft's integrity scoring used string-distance matching, including Levenshtein distance, among other methods, to decide whether a near-miss citation earns partial credit. That is a defensible engineering choice for a diagnostic, but in a filing context a citation that is almost right is wrong, and the composite should award no partial credit. The second is reporting. The composite should always be published alongside the per-dimension results, so that a low score can be diagnosed rather than merely deplored.

7. Conditions for publication

A benchmark confers a kind of definitional power: whoever publishes it defines what "good" means in a category. The draft's companion memo is candid about this, and about its commercial attraction. The Institute's view is that the power is legitimate only when constrained by the following conditions.

  1. Leave the rows where others are strong. Showing near-parity on topical recall is what makes a gap elsewhere believable.
  2. Lead with the composite, not the per-row extremes. One number, tied to the partner's actual bar, carries the argument; selected blowouts invite suspicion.
  3. Publish the rubric and make it rerunnable. A methodology a skeptic can rerun on his or her own corpus is worth more than any scoreboard.
  4. Name the integrity tie. Conceding that grounded systems do not fabricate is what earns the right to the harder claim, that not fabricating is not the same as being filing-grade.
  5. Report measurements, not targets. Published figures must be the results of runs, with dates, corpus, and configuration stated.
  6. Disclose the designer's interest. Readers are entitled to know who built the benchmark and what they build.

A disclosure about the source

The working draft adapted here was written to compare a Legawrite.AI system with three other architecture archetypes, and the Institute's founder also founded Legawrite.AI. The draft's comparative tables showed the Legawrite.AI archetype tied on fabrication with the strictly grounded archetype and leading on every other dimension, including the composite. The companion memo states plainly that those figures were hypothetical design targets rather than real measurements, and that the benchmark would carry weight only once real measurements replaced them. For that reason no figures from the draft are reproduced in this guide, and nothing here should be read as a result. The standard stands or falls on its design, not on any system's score.

8. Objections and limits

  1. Selection bias. The obvious hostile reading is that a benchmark built by a system's designer selects the tasks that system was built for. Principle 1 is exactly where that risk lives: "structural failure" is chosen by someone. The mitigations are partial and should be stated as such: publish the task-selection rationale, retain dimensions where the designer's system has no advantage, and invite third parties to contribute items.

  2. Gold answers are legal judgments. Stance, holding versus dicta, and holding scope are contested categories. Gold answers should be adjudicated by more than one annotator, with agreement rates reported. Where the law is genuinely unsettled, "the controlling line" may not exist; such items should be flagged, scored separately, or excluded. The methods discussed in Measuring Doctrinal Indeterminacy offer one way to identify them.

  3. Validity decays. An authority that is good law today may be reversed next month; the draft's own examples depend on that fact. Every gold answer must carry an as-of date, and a published run must state the date against which validity was assessed.

  4. Forum-specificity. A benchmark built on New York practice demonstrates performance on New York practice. Generalization to other forums is a claim that must be shown forum by forum, not assumed from the framework's design.

  5. Conjunctive strictness can hide diagnosis. A composite that fails a system for any single miss is the right bar for filing and the wrong tool for improvement. Hence the rule that per-dimension results are always published with it.

A checklist for evaluators

When a vendor, a researcher, or the Institute itself claims that a system is filing-grade, ask:

  • Is fabrication reported and then held constant, or is it the headline?
  • Are gold answers specified as propositions, with forum, stance, stage, and as-of date?
  • Is the composite scored conjunctively per question, and published alongside per-dimension results?
  • Do the items include department or circuit splits, subsequent history, supersession, and stage mismatches?
  • Are genuinely contested questions identified and handled separately?
  • Is the rubric public, and can the run be repeated on another corpus?
  • Are the published numbers measurements, with dates and configurations?
  • Has the designer's interest been disclosed?

A benchmark that satisfies these conditions measures what a partner checks before signing. One that does not measures something easier.

Frameworks in this piece

Terms in this piece

Revision history

1 Sep 2026Rewritten for the Institute library by Priya Raman.

How to cite

Raman, P. (2026, September 1). What a Filing-Grade Benchmark Must Measure: Dispositive recall under doctrinal and procedural constraint, and why fabrication rate is no longer the headline. Computational Law Institute. https://institute.legawrite.ai/articles/what-a-filing-grade-benchmark-must-measure

Related pieces

Follow the research

New papers, frameworks and essays. No marketing. Or use RSS.