The Long Road to Applied Computational Law
Six decades of formalizing legal reasoning, and the question of authority waiting at the end of the road

The most successful taxonomy in the history of legal technology is one that no court will allow you to cite. The West Key Number System, which originated in the nineteenth century and now runs to more than four hundred topics and nearly a hundred thousand points of law, is maintained by attorney editors and consulted by lawyers every working day. Yet its headnotes form no part of the opinions they summarize. They are not precedent, and law librarians warn that they should never be cited as law. For well over a century, the profession has navigated by a map it is forbidden to quote.
It is a small and exact image of the problem that computational law has circled for more than six decades. Law is a practice of reasons that bind. To make legal reasoning tractable to a machine, one must restate those reasons in some other form: a normalized proposition, a frame, a factor, an argument scheme, a node in a graph. The moment one does so, a question arises that no engineering can settle. By what authority does the restatement speak?
What follows is a history of the restatements. It is also an argument that the field's present "applied" turn, which promises at last to carry them out of the laboratory and into practice, inherits that question in its sharpest form.
The logician's wager
The story conventionally begins in 1957, when Layman Allen published a paper in the Yale Law Journal proposing symbolic logic as "a razor-edged tool for drafting and interpreting legal documents." His method, known as systematic pulverization, decomposed legal sentences into atomic propositions joined by five logical operators. The central insight was elegant and a little deflating: most statutory ambiguity, Allen argued, comes from confusing implication ("if P, then Q") with coimplication ("P if and only if Q"). The difference is invisible in ordinary legal English and decisive in application, because it determines whether a condition is merely sufficient or both necessary and sufficient.
Allen applied the method to portions of the Uniform Commercial Code, the Internal Revenue Code, and the Federal Rules of Civil Procedure, but never normalized a complete statute. Robert Summers's critique in 1961, that Allen had "devised an elaborate sledgehammer for cracking common-sense chestnuts," captured the profession's reluctance. In his work on normalization, Allen proposed that one should retrieve structured propositions rather than whole documents, and he named with unusual candor the question on which the whole scheme depended: how skilled must the analysts be who convert legal texts into normal form? He called it the most significant unanswered question.
Allen framed it as a question of economics, and so it was. But it is also, more quietly, a question of authority. The analyst who converts a statute into normal form decides what the statute means. Allen's wager was that ambiguity is a logical error that a trained eye can correct. The deeper difficulty, which propositional logic could not reach, was that much ambiguity is not error at all. It is vagueness, open texture, and the presence of terms that carry policy within them.
The frame and the dissent
Twenty years later, L. Thorne McCarty took a different path. His TAXMAN system, reported in 1977, did not normalize statutory language; it built deep conceptual models of corporate tax reorganization using frames and semantic networks, and performed what McCarty himself called rudimentary legal reasoning. He was explicit about the contrast that would haunt the field: the retrieval systems of his day stored full text and retrieved by keyword, whereas an extended TAXMAN would store the conceptual structure of a statute or case and retrieve by matching patterns over that structure. Its empirical base was small, fewer than ten cases, with Eisner v. Macomber (1920) as the principal test.
TAXMAN II, developed in 1980 and 1981, introduced the prototype-deformation model, McCarty's most important theoretical contribution. A prototype is a paradigmatic, undisputed instance of a legal concept; a deformation is an operation that transforms the prototype, one small step at a time, toward a contested case. McCarty reconstructed Justice Brandeis's dissent in Eisner as a chain of six such deformations running from a distribution of stock to a distribution of cash, each step a slight conceptual shift, with no principled place along the chain to draw the line of taxability.
I think this is the most philosophically charged artifact in the whole history. A machine representation, built to model legal reasoning, reproduced the argument of a dissent: that there is no line. The majority, deciding the case, had drawn one. Formalization did not discover the answer. It discovered the absence of a principled boundary, and thereby revealed that the boundary the Court announced was supplied by something other than logic. The more precisely one models legal reasoning, it turns out, the more clearly one sees where reasoning ends and decision begins.
McCarty acknowledged the two reasons TAXMAN II could not be fully implemented in 1981: the available representation languages were not expressive enough, and the prototype concept had never been adequately formalized. His Language for Legal Discourse, developed from 1989 onward, was the most ambitious representational language proposed for law; he never wrote an interpreter for the full language. In a 2022 position paper, "A Language for Legal Discourse is All You Need," he argued that a full implementation was "a feasible project today." Nearly half a century after TAXMAN, the sentence reads as both a vindication and a confession.
Statutes that compute
The clearest path to production ran through statutes. A British team translated a large part of the British Nationality Act into Prolog-like logic, so that its consequences could be derived mechanically. The project made a striking observation: the knowledge-elicitation bottleneck that crippled classical expert systems was almost entirely absent, because the law was already written down. A follow-on project on supplementary benefit legislation was equally frank that scale itself multiplied difficulties and forced ad hoc representational choices. Yet the lineage prospered. As Bench-Capon's historical retrospective records, translating legislation into executable logic became a notable commercial success through Softlaw, and, after changes, passed into Oracle's policy management product line.
The lesson was that production arrived where law was rule-like. A statute is, in a sense, a closed world: its rules and conditions can be enumerated. Case law is open: any prior decision may prove relevant, and relevance is settled by analogy, which depends on context, which is contested by adversaries.
There is a jurisprudential reason for the asymmetry, and it is not merely technical. A statute carries its own authority into the machine. Its text was enacted; the formalization is a translation of something already binding, and its legitimacy question is simply fidelity. A judicial holding has no such text. It must be constructed from an opinion, and the construction is itself an act of interpretation.
The factor and the brief
The case-based tradition took up that harder task. Kevin Ashley's HYPO, developed between 1988 and 1990, modeled adversarial argument in United States trade secret law. It represented cases along thirteen dimensions running from pro-plaintiff to pro-defendant extremes, drew on a case base of thirty-three decisions, and generated the three-ply argument every litigator knows: cite a precedent, let the opponent distinguish it, rebut the distinction. CATO, developed by Aleven and Ashley between 1995 and 1997, simplified the dimensions into twenty-six Boolean factors, each favoring one side, linked through fifty connections to sixteen abstract factors, five of them top-level legal issues, across a case base that grew to 147 decisions. It even offered a query language that translated argument constraints into retrieval constraints.
The results were real. Issue-Based Prediction, by Brüninghaus and Ashley in 2003, built on CATO's hierarchy and reached 91 percent accuracy in leave-one-out testing across 186 cases. The factor approach later traveled: to Québec landlord-tenant disputes in the JusticeBot project, and, through the Liverpool group's ANGELIC methodology, to negligence claims, the automobile exception, and Article 6 of the European Convention on Human Rights, where it matched actual Strasbourg decisions 97 percent of the time.
But the tradition was, in a phrase the working drafts use, deep but narrow. Building CATO for trade secrets alone took thousands of hours of expert modeling, and SMILE, the attempt to learn factor classifications directly from text, produced much weaker results. That was the first clear demonstration that the binding constraint was not the reasoning framework but the labor of filling it. There is a subtler limit too. The factor systems modeled argument within a doctrine; they assumed that the governing test was settled. They represented the brief, not the moment at which a court decides which test governs.
Arguments without end
From the 2000s the field turned toward formal rigor. Prakken and Sartor developed frameworks for defeasible reasoning, in which arguments conflict and priorities are themselves open to argument. The literature drew a sobering conclusion: the generic principles for resolving conflicts among rules (the specific over the general, the later over the earlier, the higher over the lower) settle only a minority of cases, because applying them depends on substantive goals and contested interpretations. Bench-Capon, Prakken and Sartor observed that analyzing legal reasoning through argument schemes, each with critical questions defining how it may be attacked, yields a taxonomy of arguments. Prakken's ASPIC+ framework of 2010 distinguished three modes of attack (undermining a premise, rebutting a conclusion, undercutting the applicability of a rule), and roughly a dozen specifically legal schemes were formalized within it. The Carneades model treated burdens of production and persuasion separately, and allowed them to shift across the stages of a dialogue. Branting defined the ratio decidendi as a justification structure and argued that "the theory under which a case is decided controls its precedential effect."
Yet no scholar has claimed an exhaustive list of valid legal reasoning moves. The best synthesis estimates somewhere between twenty-five and forty types, depending on granularity; Walton's ninety-six general schemes include perhaps twenty to twenty-five that apply to law; Eisenberg names five modes of common-law reasoning; Lamond distinguishes three models of precedent. Alexander and Sherwin argue, provocatively, that analogy cannot be formalized as a separate mode at all. The literature also concedes that coherence, the property by which a synthesis of many cases might be judged, remains hard to define. The Institute's own Twelve Bridges is a deliberately practical answer to this condition: a finite list of the doctrinal moves by which courts accept a "therefore," offered as a working discipline for practitioners rather than a proof of completeness.
A pattern emerges when these eras are set side by side. The symbolic programs assumed that human experts would do the formalizing. The argumentation programs assumed that the formalizing had already been done. What neither supplied was the middle: the reliable extraction of structure from the unstructured prose of judicial opinions, at scale, across jurisdictions.
What the market kept
Meanwhile, what actually scaled was simpler: editorial taxonomies such as the Key Number System, citators that classify how later courts have treated a case without classifying what it holds, and standards for describing legal matters and documents. The statistical turn, roughly from 1995 to 2015, routed around the formalization problem rather than solving it; a model trained on legal text could predict outcomes but could not explain itself in terms a lawyer could audit or a court could cite. The present generation of retrieval-assisted research tools inherits both habits; a 2025 empirical study of major legal research products found that hallucinations and inaccuracies persist, distinguished reasoning errors from retrieval failures, and observed that retrieval itself often requires legal issue-spotting.
Here the opening paradox returns in a more general form. What scaled stood, in one way or another, outside law's economy of reasons. The editorial map is authoritative by no one's enactment. The statistical prediction gives no reasons at all.
The applied turn
The Institute's working drafts argue that the ground has now shifted. Their verdict on the history is generous: no complete, computationally tractable taxonomy of holdings, pleading defects, or synthesis validation was ever built, but the fragments produced between 1957 and 2025 collectively map roughly 60 to 70 percent of what a finished architecture would require. The obstacle was never theoretical impossibility. It was the knowledge-engineering bottleneck.
Three developments, on this account, have changed the calculus. The Harvard Caselaw Access Project offers some 6.7 million decisions, where HYPO had thirty-three. Machine-assisted annotation lowers the cost per example by an order of magnitude while holding accuracy near an F1 of 0.80. And the CaseHOLD insight, that citing opinions already contain descriptions of the holdings they rely upon, provides a way to harvest holding statements at scale without annotating them afresh. The drafts conclude that the theoretical work is substantially done, and that the question is no longer whether the formal structures exist but whether anyone will fund the labor of filling them.
One of the drafts goes further. Its thesis is that the fifty-year barrier was engineering: language models now extract holdings "well enough, at scale," graph databases now store and traverse structured legal knowledge, and the research of the preceding decades "was not wrong. It was waiting." It locates the new bottleneck elsewhere: "trust, integration, and scale." That draft was written from inside one such effort, Legawrite.AI, founded by the director of this Institute, and it describes that company's own system as having carried these ideas into production. It is right to record the claim and right to treat it as a claim. The Institute's fuller scholarly case for proposition-level representation at corpus scale is made in The Promise Fulfilled.
Whatever one concludes about any particular system, the fragments may, for the first time, be filled. That is precisely when the old question of authority ceases to be academic.
Four questions the turn must answer
Who authors the holding? Branting's observation, that the theory under which a case is decided controls its precedential effect, cuts deeper than it first appears. To extract "the holding" of an opinion, or to label one case as having established a rule and another as merely applying it, is to choose a theory of the case. When a court does this, it is adjudication. When an editor does it, it is commentary. A machine-extracted holding belongs, by its nature, with the headnote rather than the opinion. The danger is not chiefly that it will be wrong. The danger is that it will be right often enough to be relied upon as though it were the law: the Key Number paradox, at scale and at speed.
Who draws the line where reasons run out? McCarty's reconstruction of the Eisner dissent showed that formalization can locate the place where a chain of legal reasons stops compelling the next step. At that place, the legitimacy of a conclusion comes from an institution entitled to decide, not from any model of the reasoning. A system that quietly draws the line there has usurped a function; a system that marks the place and says "judgment is required here" has honored a division of labor. The applied draft, to its credit, puts its ambition in the second form: to reason about the parts that can be reasoned about, and to flag the parts that require judgment. The Institute's work on measuring doctrinal indeterminacy is an attempt to make that flag precise.
Can the reasons be given? Legal authority is inseparable from reason-giving; the statistical turn gave this up. The applied turn's promise is structure that can be inspected, a chain from conclusion to node to opinion that a lawyer can follow and, if necessary, break. But McCarty warned in 2022 that prompt engineering remains "ad hoc empiricism" rather than principled representation. If the layer that extracts the structure is itself unprincipled, an inspectable graph may rest on an uninspectable foundation. The only adequate answer is fidelity to text: every node traceable to the words of a court.
Who governs the taxonomy? Allen's unanswered question about the skill of the analysts returns here as a question of governance. The working drafts observe that the bottleneck is rarely the existence of a taxonomy; it is the cost and governance of maintaining one across time, jurisdictions, and domains. The enumeration of valid reasoning moves is incomplete, and coherence is undefined. To validate a lawyer's synthesis against a list that no one claims is complete is to exercise a form of authority over legal reasoning itself. Such authority is not illegitimate. But it should be exercised openly, with its criteria published and its incompleteness confessed.
Coda
Picture the rooms in which much of this history was made: a few researchers, a few machines, a case base of thirty-three trade secret decisions, and the audacious question whether law could be computed. After six decades the answer seems to be that a great deal of it can, and that the parts which cannot are exactly the parts where law's authority lives: the drawing of lines that reasons do not fix, and the choice among theories that the materials permit.
The road to applied computational law was long because the labor was immense. Its destination is not an answer but a question. It is no longer whether machines can represent legal reasoning. It is by whose leave, and under what discipline of reasons, their representations will be believed.
Frameworks in this piece
Terms in this piece
Revision history
| 23 Jul 2026 | Rewritten for the Institute library by Julian Fenwick. |
How to cite
Fenwick, J. (2026, July 23). The Long Road to Applied Computational Law: Six decades of formalizing legal reasoning, and the question of authority waiting at the end of the road. Computational Law Institute. https://institute.legawrite.ai/articles/the-long-road-to-applied-computational-law
Related pieces
New papers, frameworks and essays. No marketing. Or use RSS.


