The Docket Test
How to Evaluate Legal AI for Dispositive Motions in One Afternoon

Bottom line: don't buy a legal research tool on a demo. Buy it on a motion your own court already decided. The answer key exists. It's on the docket.
Every vendor will show you a clean screen, fast answers, and citations that check out. None of that tells you what you need to know, because motions aren't lost on the citations that are there. They're lost on the controlling case nobody pulled, the element nobody briefed, and the case on the other side that showed up for the first time in the opposition. A demo can't show you what's missing. A decided motion can.
Here's why. A decided summary judgment motion or motion to dismiss gives you four documents: the moving brief, an opposition written by someone whose job was to find everything the movant missed, a reply, and an order where a judge says what actually mattered. That's an answer key built by two motivated adversaries and graded by a judge. You can't buy a better benchmark. It costs a few dollars in PACER fees, or nothing if the filings are on CourtListener.
This is the test. One motion, one tool, one afternoon.
Step 1: Pick the motion (30 minutes)
You want one that looks like your actual work.
- Your court, your kind of case. A test in someone else's jurisdiction on someone else's claims tells you about their practice, not yours.
- Recent. Ideally decided after the tool's training data was collected, so it can't have memorized the answer. Unpublished orders are good for this.
- A denial or partial denial, if you can find one. Orders granting summary judgment often decide on one element and skip the rest. An order denying the motion has to identify the genuine dispute, which gives you more to score against.
- The full set. Motion, opposition, reply, and order. Without the opposition you lose half the answer key.
Step 2: Run the tool (90 minutes)
Give the tool only what the moving party had on the day it filed:
- the complaint (and answer, if one was filed);
- the governing law;
- the question the motion posed;
- for summary judgment, the record excerpts the parties had.
Do not give it the opposition or the order. Then ask it to do the research and outline the motion. If you'd rather test the other side, give it the moving papers and ask for the opposition.
Save everything it returns.
Step 3: Score it (60 minutes)
Lay the tool's output next to the opposition and the order. Fill in this sheet.
| # | What you're scoring | How to count it | Why it matters |
|---|---|---|---|
| 1 | Controlling authority omitted | Of the authorities the order relied on, how many did the tool never surface? | This is the headline number. A missed controlling case loses motions. |
| 2 | Adverse authority surfaced | Of the authorities the opposition cited as controlling that the movant never cited, how many did the tool put in front of you, whether cited, distinguished, or flagged? | This is the case you'd otherwise meet for the first time in the opposition. |
| 3 | Posture mismatches | How many citations were decided under a different stage, standard, burden, or forum, and cited for something that depended on it? | Real cases, wrong motion. No citator flags these. |
| 4 | Element coverage | Of the elements the order analyzed or the opposition briefed, how many did the tool address? | A missed element can defeat the motion by itself. |
| 5 | Fabricated or misgrounded citations | Cases that don't exist, or real cases cited for something they don't say | Still table stakes. Count them. |
That's the whole test.
How to read the results
Three rules. They're what keep this from turning into another vendor slide.
Don't average the five lines into one score. A tool that's perfect on line 5 and bad on line 1 is dangerous, because it looks trustworthy while missing the case that matters. Report the five numbers side by side. If one tool wins on some lines and loses on others, say which.
One motion is an anecdote. Three is a decision. Run at least three motions before you sign anything. Include one where your side lost.
Run the same motions on every tool you're considering. Any quirk in the motions you picked affects every tool equally, so the side-by-side comparison holds up even when the absolute numbers don't.
Two honest limits. First, courts often decide on narrow grounds, so the order won't list every authority a complete brief needed. Scoring against the order plus the opposition covers most of that gap. Second, briefs are advocacy, and each side left out things that hurt it. Scoring against both sides' briefs is how you correct for that. Neither limit changes the comparison between tools.
What this costs, in plain numbers
One afternoon of a senior associate's time per tool, for the first motion. The second and third go faster once you have the score sheet.
Compare that to the cost of one summary judgment motion lost on a case nobody pulled. That's not a research budget line. That's the case.
If a vendor won't let you run this during a trial period, that tells you what you need to know.
Here's what I want from you: five questions for every vendor
Ask these in writing. Keep the answers.
- What is your measured rate of missed controlling authority, and who measured it? I want the method, the sample, and an evaluator who isn't you. "Our citations are verified" doesn't answer this question.
- How does your system know what stage and standard a holding was decided under, and can I filter on it? If the answer is "the model figures it out," the answer is no.
- When I ask for cases against my position, how does the system know which way a holding cuts? Same answer, same problem.
- Can I see the record for any answer? What was searched, what was retrieved, what was discarded, and why. I want to export it and put it in the file.
- What does your system do when it can't find something? Does it tell me, element by element, or does it write around the gap? And what are your refusal and incompleteness rates?
Any vendor that answers all five clearly is worth an afternoon. Any vendor that won't isn't worth a contract.
Disclosure: this test is published by Legawrite.AI. We build our product to be scored this way, and we'd rather you run the test on us than take our word for it.
Sources and method: The scoring approach adapts the docket-based evaluation design in which a decided motion's briefs and order serve as adversarially produced ground truth. Relevant authorities: Fed. R. Civ. P. 12(b)(6), 56(a), 56(c)(3); Celotex Corp. v. Catrett, 477 U.S. 317 (1986); Model Rules of Prof'l Conduct R. 3.3(a)(2); ABA Formal Opinion 512, Generative Artificial Intelligence Tools (2024).
Terms in this piece
Revision history
| 26 Jun 2026 | First published in the Institute library. |
How to cite
Brennan, T. (2026, June 26). The Docket Test: How to Evaluate Legal AI for Dispositive Motions in One Afternoon. Computational Law Institute. https://institute.legawrite.ai/articles/the-docket-test
Related pieces
New papers, frameworks and essays. No marketing. Or use RSS.


