What Filing-Grade Means
The Institute's working standard for machine-assisted legal work product: nine criteria a brief or memorandum must meet before it is fit to file, and how each will be measured in the open.
Updated 1 September 2026
The Institute's charter commits it to define what "filing-grade" means for machine-assisted legal work and to measure it in the open: not merely whether an AI system fabricates citations, but whether it finds the controlling authority, surfaces the adverse precedent, confirms that a holding is still good law, and shows the exact reasoning path from precedent to conclusion. This page makes it a working standard.
Filing-grade is a property of a work product, not of a system: a brief meets the standard or it does not, whoever or whatever drafted it. The standard is written so that any system can fail it in public, including systems built at Legawrite.AI, the company our founder also founded. A grounded system can hold fabrication at zero and still hand over a real, accurately quoted case that has been reversed, was decided at the wrong stage, or is simply the wrong case for the job. Each criterion below names a failure that passes a citation check.
The criteria
A work product is filing-grade when it meets all nine. They are not weighted or averaged: a draft that meets eight has a defect, not a high score.
-
Controlling authority found. For every issue, the authority that decides it in this forum is in the draft. No checker can see its absence: a verifier reads the document, so a controlling case that was never retrieved never enters its input. Research therefore runs top-down, from the forum's court of last resort to the appellate court that binds the trial court, before any persuasive authority. Developed in The Case You Never Pulled and the Pre-Filing Completeness Protocol.
-
Adverse authority surfaced. The controlling authority against the position is found and put in front of the lawyer, whether the draft then cites, distinguishes or argues around it. Model Rule 3.3(a)(2) requires disclosure of directly adverse controlling authority known to the lawyer and not disclosed by opposing counsel; a process that never looks for it leaves that duty to chance. Direction is a property of a holding, not of a case, so every load-bearing rule needs three further searches: its exceptions, the cure paths open to the other side, and the cases that distinguished it. Developed in Read It Like Opposing Counsel and the Counter-Model Builder.
-
Good-law status confirmed at the proposition level. A lawyer cites a proposition and offers a case as its container. One opinion can be overruled on one holding, distinguished on a second and followed on a third, and a single case-level flag cannot say so. A doctrine can also survive formally while its practical direction shifts, as the Auer deference line did once Kisor v. Wilkie imposed prerequisites on it. Nor is a citator's flag settled fact: in one study of 357 citing relationships that at least one of three citators labeled negative, all three agreed on negative treatment in only 53. Developed in the Proposition-Usability Model, the Case Treatment Taxonomy and Good Law for What?
-
Procedural posture and forum fit. Every holding was announced under a standard, applied to a defined body of material, with a burden on a particular party, in a particular forum. A summary judgment opinion saying the plaintiff "offered no evidence" of an element describes a record and does not belong in a motion confined to the pleadings. Twombly's plausibility standard does not decide a California demurrer, which turns on the state's fact-pleading statute. A citator flags neither, because nothing in the later history changed. The test is whether the proposition survives translation into this motion's standard, record, burden and forum. Developed in Right Law, Wrong Stage, Twombly Doesn't Live Here and the Posture Mismatch Taxonomy.
-
Every element addressed. Each element of each claim and defense at issue is analyzed, with the party bearing the burden identified. Element omission sits upstream of everything else: retrieval, generation and verification are all invoked by questions, so an element nobody asks about produces no query, no citation and no error for any tool to catch. Dispositive motions make this checkable, because the governing substantive law supplies the element list before research begins. Developed in the Pre-Filing Completeness Protocol.
-
Record support cited. Every factual assertion bearing on a contested element carries a pinpoint citation to the record. At summary judgment, Rule 56(c)(1)(A) requires "citing to particular parts of materials in the record," and Rule 56(c)(3) provides that the court "need consider only the cited materials." Evidence left uncited may go unweighed, so an uncited deposition passage can decide whether a triable issue exists. Developed in the Pre-Filing Completeness Protocol and Research the Carrier Won't Pay For.
-
Rejected arguments checked. Before an argument is made, the lawyer knows whether courts in the forum have already considered and refused it. A rejected argument is a party's contention, described by a court in order to refuse it: not a holding, rarely a headnote, and invisible to citators, because nothing happened to the cited authority. Finding it takes a different search from finding the rule: state the argument as one proposition (this rule, applied to these facts, yields this conclusion) and search for it with rejection language in the forum. Developed in Good Law for What?, The Promise Fulfilled and the lexicon entry on negative space.
-
Reasoning path from precedent to conclusion shown. Each conclusion can be traced to its authorities and to the mode of inference that carries it from them. Real authorities joined by an invalid inference yield a conclusion no citation database can catch, because every citation checks out. The Twelve Bridges framework names the legitimate doctrinal mechanisms for deriving a conclusion from precedent, together with the fault patterns that mark an invalid one. The path belongs in a record that travels with the work: what was asked, searched, retrieved, rejected and why, and what a human reviewed. Developed in the Twelve Bridges, the First Law of the Four Laws and the lexicon entry on the derivation record.
-
Calibrated confidence and the ability to abstain. The work product states how sure it is and reports, element by element, what it could not find. The Zeroth of the Four Laws outranks the rest: never present an output with more confidence than can be defended. Abstention is part of the standard, but it must be counted, because a system can lower its error rate by declining to answer; refusal and incompleteness rates are reported next to error rates. The Institute's own instruments follow the rule: the posture-dispersion measure in Ariadne's Thread issues no reading when the signal does not clear the noise floor, and in our blind eDiscovery review a failed calibration gate sent every responsive determination to human review. Developed in the Four Laws, posture dispersion Z and We Ran a Blind eDiscovery Review.
How it will be measured
A standard nobody can check is an opinion. These criteria will be measured by an open, rerunnable benchmark with a public rubric, fixed before any system is run and published with the results, so that anyone can rerun it.
- Completeness and omission are graded alongside accuracy. Beyond the citations a document contains, the benchmark scores what it should have contained: controlling authority omitted, adverse authority surfaced, element coverage and record support, next to posture mismatches and fabricated or misgrounded citations. Fabrication is reported and then set aside, so that comparisons among systems that do not fabricate turn on what actually differs.
- Decided motions serve as answer keys. As The Docket Test sets out, a decided dispositive motion comes with a moving brief, an opposition written to find everything the movant missed, a reply, and an order saying what mattered: ground truth produced by two motivated adversaries and adjudicated by a judge. Because courts often decide on narrow grounds, scoring runs against the order and the opposition together, and motions decided after a system's training data was collected guard against memorized answers.
- Results are reported side by side, never averaged. A tool that is perfect on fabrication and weak on controlling authority looks trustworthy while missing the case that matters; an average hides that. Where a single figure helps, it is a conjunction: an answer counts only if it clears every constraint at once, as in What a Filing-Grade Benchmark Must Measure.
Benchmark materials will appear on Open Source and Data as they are released. Failures observed in court are logged in the Citation Failure Record.
Version
Version 0.1, 1 September 2026. It is open for comment. Criteria may be split, merged, sharpened or retired as the benchmark produces evidence, and each change will be dated and noted here. The most useful comments point to a decided motion, a court order or a published study that the current text gets wrong.