AXL-WP-09 · COMPANION WRITTEN FOR HUMANS
AXL-WP-09 THE PLAIN VERSION ~4 MIN

Pay people to find the flaw.

Grant review can't tell good from great, and AI just made every proposal look great. Here's a review system that finds out whether it was right.

c/o IMRAN HAFIZ · "WRITTEN FOR HUMANS, NOT REVIEWERS" · AUSTIN, TX

The plain version: grant funders pick winners using signals that have stopped meaning anything. A polished proposal used to suggest serious effort, and AI now produces polish for free. Expert review was already struggling: experienced reviewers scoring the same funded proposals barely agreed with each other. This paper designs a review process that pays people to find flaws, and lets a funder finally learn whether its picks were right.

Here's a number that stuck with me. In one Australian funding round, researchers spent about 550 working years writing 632 proposals, and the time someone put in did not predict whether they got funded. That was measured in 2013, long before AI. The cost of writing a proposal never measured the idea inside it. It measured who had the time and the support to write it. Now that cost is close to zero, and funders are left sorting on noise.

The usual answer is that expert reviewers read past the polish to the substance. But when researchers had experienced NIH reviewers rescore applications that had already been funded, each reviewer was consistent with themselves and almost nobody agreed with anybody else. The problem isn't laziness. Nothing in the system pays anyone to find a flaw. A careful "this holds up" looks exactly like skimming. And no reviewer is ever told whether they were right, so the process never learns.

The Crucible changes the incentives. Anyone can submit a proposal, and anyone can attack one. A panel drawn at random from registered experts judges the proposals and the attacks, and is paid using a scoring method built so that an informed minority can beat a confident majority. Nobody bets any money. Confirming that a proposal is sound pays the same as tearing one down, so empty theatrics don't pay. The proposals that survive are funded by a weighted lottery, with the odds published in advance.

The lottery is the part that sounds strange, and it's the part already in use. Four research funders, in Switzerland, Germany, Austria and New Zealand, already fund some grants by lottery among proposals their reviewers couldn't separate. They did it for fairness. None of them uses what the lottery makes possible: random selection is an experiment. Compare what happens to the funded and the unfunded proposals, and a funder can finally measure how accurate its reviews were.

I'll be straight about the status. Nothing here has been run yet. It's a specification, published so people can break it. Its weakest point is time: outcomes take three to five years to show up, so the first real answer is years away.

[ WHY_THIS_MATTERS ]

Grant money is how a society decides which ideas get a chance. If the filter can't tell good from great, and nobody ever checks, we end up funding whoever writes the most convincing document. AI makes that a little worse every month.

[ THE_PREDICTIONS ]

The paper commits to withdrawing its claims if any of these happen:

  • The scoring method can't pick out informed minority views on panels of people from different backgrounds.
  • Flashy, low-substance critiques are systematically paid more than substantive ones.
  • Confirming that a proposal holds up pays materially less than refuting one.
  • A funder ever loses more than the fixed ceiling set before the round began.
  • The first comparison against lottery outcomes shows the panel and the market doing no better than a simple unweighted baseline.

That last one is the one I'd least like to lose. The full design, the math and the open problems are in the paper.

[ READ_THE_FULL_PAPER: AXL-WP-09 -> ]

Does your organization need thinking like this?
Book a meeting · imran@auxlabs.co