Lawyers are right to extend no default trust to a machine, and a published claim about a machine deserves no more. A benchmark result is worth what a skeptic can rerun. A defensibility claim is worth what an opposing party, a special master or a court-appointed expert can reproduce. That is why the Institute publishes its instruments and not only its conclusions. As the report on our blind eDiscovery review put it: "Openness is not marketing. It is the mechanism of trust."
Openness has to include what did not go well. When Legawrite.AI released isResponsive, the execution logs and failure analysis went out with the code, and the report on the run disclosed the two validation gates it failed. The same rule governs this page. Each item states its status plainly (released, in development or planned) and lists only what we can confirm. Where a repository address has not been verified in our sources, we give the install command and no link rather than a guessed one. Written protocols and label sets count as released once their full text is published, because a method anyone can apply by hand is already open.
codereleased
isResponsive (legawrite-tar)
An open TAR 3.0 engine for court-defensible responsiveness review, released by Legawrite.AI with its pipeline architecture, execution logs, metrics, failure analysis and reproducibility records. It classifies documents against per-request rubrics locked by cryptographic hash, cites source text with a calibrated confidence for every decision, and routes privilege flags to attorneys.
pip install legawrite-tar
benchmarkin development
ADR-Hard: Adversarial Doctrinal Retrieval, Hard Tier
A filing-grade benchmark for legal research systems that measures dispositive recall under doctrinal and procedural constraint, rather than topical relevance or fabrication alone. Its rubric credits an answer only when it states the right proposition with the right stance, at the right procedural stage, with confirmed good-law status, and currently citable in the forum.
protocolreleased
A written protocol for evaluating a legal AI tool in one afternoon against a motion the lawyer's own court has already decided, using the moving brief, opposition, reply and order as an answer key built by two adversaries and graded by a judge. It reports five numbers side by side and never averages them: controlling authority omitted, adverse authority surfaced, posture mismatches, element coverage, and fabricated or misgrounded citations.
protocolreleased
A written pre-filing protocol that builds a matrix of every element of every claim at issue against five checks: governing standard, controlling authority, record support, adverse authority, and rejected arguments. No cell may be left blank; each ends in a citation, a logged search that found nothing, or a one-line reason it does not apply.
datasetin development
A public, source-cited record of court decisions on AI citation failures in filings, classified by failure type beyond fabrication. The first release is a small seed set drawn from the Institute's corpus, not a census, and every entry cites a court order or a published secondary source.
datasetreleased
A written label set for classifying how a later opinion treats an earlier one, applied to specific holdings rather than whole cases and described along three dimensions: scope, severity and mechanism. Its labels run from negative (such as overruled, abrogated, reversed, vacated and superseded by statute) through cautionary and positive to neutral, and it resolves doubt toward flagging rather than away from it.