AXL-WP-07 · COMPANION WRITTEN FOR HUMANS
AXL-WP-07 THE PLAIN VERSION ~5 MIN

The courts are fast. That's the problem.

American courts grade themselves on speed. AI is about to show us exactly why that was a mistake.

c/o IMRAN HAFIZ · "WRITTEN FOR HUMANS, NOT REVIEWERS" · AUSTIN, TX

The plain version: courts measure success by how quickly they clear cases. By that measure, a court can look excellent while most of the people in it lose automatically, without anyone ever checking the facts. AI-generated filings are about to flood that system, and the speedometer can't see the flood.

I studied public policy at Duke, and one of the first lessons that field beats into you is that institutions become what they measure. Give a bureaucracy a number and it will optimize the number, shedding everything the number doesn't capture. It's true of schools, police departments, and hospitals. This paper is about what it did to the courts.

The number courts live by is the clearance rate: cases resolved divided by cases filed. It sounds reasonable. Here's what it hides. In American consumer-debt cases, more than seven in ten defendants lose by default, meaning no hearing, no examination of evidence, nobody checked whether the claim was even accurate. Fewer than one in ten had a lawyer. Every one of those defaults clears a case. The dashboard reads efficient. Whether anything resembling justice occurred is, quite literally, not measured.

Now add AI. Anyone can generate a hundred pages of fluent, official-sounding legal filing in an afternoon, and people representing themselves increasingly do. Clerks and judges are drowning in text that looks like law but hasn't been checked by anyone, including the person who filed it. Under a speed metric, the rational institutional response to that flood is the terrifying one: process it faster. More defaults, more rubber stamps, more clearance. The metric rewards exactly the collapse it should be catching.

The paper's proposal borrows a trick from economic history that I find genuinely beautiful. For decades, economists measured national well-being by GDP, until some of them noticed that in certain "growing" economies, people were getting shorter. Average height quietly recorded the nutrition and childhood hardship that the money numbers missed. Bodies kept a ledger the dashboard didn't. The paper asks: what is the courts' equivalent of height? What would we measure if we wanted the truth to show up whether or not the institution wanted to report it? It proposes a specific answer, an index built so that a court can't score well by defaulting people faster, and shows the math for why the usual way of averaging performance hides exactly the failures that matter.

I'll be straight about the status: the instrument is a proposal, not a finding. It hasn't been fielded yet. But the problem is not hypothetical, it is compounding monthly, and whoever builds the credit score for machine-era court filings will shape how an entire branch of government meets this decade.

[ WHY_THIS_MATTERS ]

When a case is resolved against someone who never got a hearing, the error does not stay in the courthouse. It follows people out the door, and it compounds into exactly the collapse of institutional trust this whole site is about, at the one institution meant to be the last resort.

[ THE_PREDICTIONS ]

The paper names four ways the instrument fails, in advance:

  • If courts score better on the index while fewer cases reach their merits, I have built a faster injustice machine, and it has failed on its own terms.
  • If hallucinated citations do not decline after deployment, the veracity check is not working.
  • If trained scorers cannot agree on the rubrics, this is not measurable in the field.
  • If the index moves while reality does not, it got gamed, and the gaming-resistance did not survive contact.

The framework, the math, and the falsification conditions are in the paper.

[ READ_THE_FULL_PAPER: AXL-WP-07 -> ]