AXL-WP-09 · v2.1 AUX LABS RESEARCH
AXL-WP-09 WORKING PAPER · v2.1 AUX LABS LLC 2,818 WORDS · ~12 MIN READ

The Crucible: Paying People to Attack Grant Proposals

ABSTRACT

Grantmakers sort applications using a signal that has stopped carrying information. Proposal polish was never a good measure of proposal quality (Australian researchers spent 550 working years on a single NHMRC round, and the time invested did not predict success), and generative models have now driven its cost to zero. The fallback, expert peer review, was already failing the same test: experienced NIH reviewers rescoring previously funded applications showed essentially no agreement with one another, and the authors of that study concluded the process is not built to tell good proposals from great ones. The common cause is that nothing in grant review pays anyone to find a flaw, and nobody ever learns whether a review was right. This paper describes the Crucible, a mechanism that pays experts to attack proposals, prices which ones survive, and allocates by weighted lottery among the survivors. Four funders (the Swiss National Science Foundation, the Volkswagen Foundation, the Austrian Science Fund and the Health Research Council of New Zealand) already allocate part of their budgets by lottery among applications peer review cannot separate. Each of them is running a randomised experiment and discarding the result. The argument of this paper is that the randomisation they adopted for fairness is also an identification strategy, and that adding a paid adversarial front end turns a fairness concession into the only grant process that can measure its own accuracy.

KEYWORDS: GRANT PEER REVIEW, INFORMATION ELICITATION WITHOUT VERIFICATION, SURPRISING POPULARITY, MARKET SCORING RULES, PARTIAL LOTTERIES, RESEARCH FUNDING

CITE AS: HAFIZ, I. (2026). The Crucible. AUX LABS WORKING PAPER AXL-WP-09. AUXLABS.CO

CONTACT: imran@auxlabs.co

1. The screen is gone

Writing a fundable proposal used to take weeks. In one Australian National Health and Medical Research Council round, 632 proposals consumed about 550 working years of researchers’ time, roughly 34 working days each, against a success rate of 21 per cent. Four centuries of effort produced nothing.

The detail worth holding onto is this: the researchers who spent more time on a proposal were no more likely to be funded. That was measured in 2013, long before anyone could generate a literature review on request. The expense of writing a proposal was never a measure of the quality of the idea inside it. It was a measure of who had the time, the training and the administrative support to produce the artefact. Funders sorted on it anyway, because it was the only thing in front of them.

Generative models have now taken that expense to zero. The structure, the citations, the register, the hedging, the limitations section that carefully names only the safe limitations: all of it is available in minutes to anyone. What has not changed by a single dollar is the cost of working out whether the argument underneath is any good.

So a signal that was already weak has gone to noise, and the volume on top of it is rising. Applications go up. Discriminability goes down. Nothing in a program officer’s reporting registers that anything has happened, because “we could not tell” is not a field in the system.

Spending longer on a proposal never made it more likely to be funded. The filter was never measuring what everyone believed it measured.

2. Peer review was already failing this test

The standard answer is that polish does not matter because expert reviewers read the substance. The evidence does not support it.

Pier and colleagues gave experienced NIH reviewers a set of applications that had already been funded, some outright and some after revision, and had them score the applications again under realistic study-section conditions. Reviewers were internally consistent: each applied their own standard steadily. Across reviewers there was almost no agreement at all. The same application drew scores so different that, as the authors put it, reviewers might have been reading different documents. The gap was not in spotting weaknesses; it was in how any given reviewer converted weaknesses into a number.

Their conclusion was not that peer review is corrupt or lazy. It was that the process is not designed to discriminate between good and excellent applications, and that after the weak submissions are cleared, the remaining ranking carries much less information than the decisions built on it assume.

Why would it? Look at what the system actually rewards.

Nobody is paid to find a flaw. Review is unpaid or nominally paid, recruited on professional obligation, and performed in the gaps of a working week.

A reviewer who reads carefully, finds nothing wrong, and says so has produced an outcome indistinguishable from a reviewer who skimmed. Thoroughness that confirms is invisible.

The work is private. No critique is on the record, nothing is contested in the open, and no reviewer’s judgment is ever attached to their name in a way that can be checked later.

And at the root of all three: no reviewer is ever told whether they were right. There is no point at which a score is compared to an outcome and a reviewer learns anything. A system with no feedback loop does not improve, and after enough cycles it stops being an evaluation and becomes a ritual that produces a number.

3. The Crucible

The Crucible is a room built to attack proposals on purpose, and to pay the people who attack them well.

A funder publishes a question and underwrites one round. Anyone may submit a proposal. Anyone may attack one. Nobody bets: there is no counterparty, no deposit, no account to open, nothing a participant can stake and lose. One cheque goes in from the funder and is spent on getting an answer.

Three rooms, and each one informs the next without deciding for it.

The Room: evaluation. A panel is drawn at random from a registered pool of experts and asked two questions about each proposal and each attack on it: what do you think, and what do you think the rest of the panel will say. Payment follows the judgment that turns out to be more widely held than the panel predicted: the surprisingly popular answer (Prelec, Seung & McCoy 2017), a method built for exactly the case where the majority is wrong and an informed minority is right.

Three things follow, and they are the reason this is the scoring rule and not another.

The author gets no vote on the value of the attack against them, so suppression and log-rolling are removed structurally rather than policed.

An attack that is exciting and empty is precisely the attack everyone predicted would land. Its surprise is near zero and it earns accordingly. The mechanism will not pay for theatre, and it will not pay for consensus either.

A reviewer who examines a proposal, finds it sound, and says so against an expectation of demolition is surprisingly popular and is paid as such. Half of real review is confirming that something holds up. No existing system pays a cent for it; this one pays the same as it pays for a kill.

The Board: pricing. A bounded-loss market maker prices one proposition per surviving proposal: will a fresh independent panel, drawn later and paired with an audit, judge this sound? Not “was it adopted”, because a single interested official can deliver an adoption, and a settlement condition one party can satisfy is not a settlement condition. A panel that has not been drawn yet cannot be captured.

The funder underwrites this. For a logarithmic market scoring rule over n proposals with liquidity parameter b, worst-case exposure is b · ln(n): with nine proposals and a $50,000 ceiling, b ≈ $22,760. That number is fixed before the first submission and holds against any sequence of trades. It is a property of the cost function, not a budget someone has to be disciplined about.

The Draw: allocation. Funding goes by weighted lottery among everything above the threshold, probabilities published in advance.

Figure 1. The Crucible’s three rooms: evaluation by a random paid panel, pricing by a bounded-loss market, allocation by a published weighted lottery.
FIG. 01 The Crucible’s three rooms: evaluation by a random paid panel, pricing by a bounded-loss market, allocation by a published weighted lottery.

This is sometimes read as the strange part. It is the part that already exists in the field, and §5 is about that.

What the funder buys

A ranked set of proposals in which every serious objection has already been raised in public, answered or conceded on the record, and priced. A paid, open adversarial process in place of three unpaid reviewers working alone. A hard ceiling on exposure known on day one. And, eventually, an error rate on its own judgments, which no grant process currently produces.

What a reviewer gets

Payment in weeks, not thanks in a year. Standing that is public and attached to their name. And a reason to write down the thing that practitioners always know and nobody puts in a review: not is the science sound but this will not survive contact with the people who have to run it, and here is the specific reason. That critique is worth money here and worth nothing anywhere else.

Why “anti-Kalshi”

The machinery is borrowed from prediction markets and pointed the other way. A retail event market harvests the attention of people who can least afford it in order to price questions nobody needed priced. This harvests the attention of people who have expertise, to price proposals that someone is about to spend real money on, with no wager anywhere in the structure. Same instrument. Opposite extraction vector.

4. Against the current process

Panel review today The Crucible
Who is paid to find flaws Nobody Anyone who finds one
A correct “this is sound” Indistinguishable from skimming Paid like a refutation
Where critique happens Private, unattributed Public, on the record, named
Reviewer disagreement Averaged away into a score Priced, because it is the signal
Does the funder learn if review was right Never Yes, from §5
Cost ceiling Soft; scales with reviewer hours b · ln(n), fixed at funding
At triple the volume Queue becomes the filter Open entry; attacks scale with payment

Four of those rows are consequences of one design choice, which is paying for the attack. The rest follow from putting it in the open.

The volume row is the one that decides this. Panel review scales linearly in expert hours, and expert hours are the one input that cannot be increased. The people qualified to review are the same people writing the applications, and they are already at the limit. Every additional application makes the queue longer and the per-application attention smaller. The Crucible scales the other way: more submissions mean more surface to attack, more attacks mean more payment claimed against a fixed subsidy, and the open pool is not capped by whoever a program officer happens to know.

The honest form of the comparison: current panel review is cheaper per application today and will stay that way. What it cannot do at any price is tell you whether it was right. The Crucible costs more per round and produces an auditable number. A funder who does not want that number should not buy this.

5. The experiment is already running

The weighted lottery in §3 is not a thought experiment imposed on funders by a theorist. It is current practice at four research funders, each randomising among applications that review could not separate.

The Swiss National Science Foundation piloted randomisation in its Postdoc.Mobility scheme from 2019, applying it to a small share of applications (about four per cent), with a Bayesian ranking defining which enter the random pool. The Volkswagen Foundation’s Experiment! programme reported increased diversity in the projects funded and more risk-taking in what was submitted. The Austrian Science Fund’s 1000 Ideas runs double-blind review followed by a random draw at the board meeting, and has recorded no applicant complaints. The Health Research Council of New Zealand’s Explorer Grants have seen acceptability rise among both applicants and assessors.

Pier and colleagues reached the same place from the evidence: having shown that reviewers cannot separate good from excellent, they proposed clearing the weak applications and then funding the remainder at random.

So the step that looks radical has four implementations and an endorsement from the people who measured the problem. Every one of them adopted it for the same reason: fairness, when the ranking is known to be arbitrary at the margin.

And every one of them then throws the result away.

A weighted lottery among comparable applications is random assignment. Some proposals are funded and others, above the same bar, differing only by the draw, are not. That is the structure of an experiment, and it is sitting unused inside four funding programmes right now. Nobody is collecting outcomes on the unfunded arm. Nobody is scoring the panel that set the threshold against what the lottery later revealed. The randomisation is treated as an apology for the limits of review rather than as the instrument that could measure them.

This is the claim the paper exists to make. The randomisation these funders adopted for fairness is also an identification strategy, and reading it as one converts the weakest moment in grant-making into the only place it can learn anything. Panel scores and market prices, recorded before the draw, are predictions. The funded and unfunded arms eventually produce outcomes. Compare them and the unverifiable front end acquires a published error rate.

The Crucible is what makes that comparison worth running. A lottery over scores nobody is paid to contest measures the accuracy of a process that was barely trying. A lottery over proposals that have survived paid public attack measures something a funder would actually want to know.

6. What could kill it

The loop does not close for years. Outcomes on funded and unfunded arms arrive three to five years out. Until then the error rate is an intention, and early rounds run on an instrument with no track record. This cannot be engineered away and is stated here rather than buried.

Shared priors. Surprising-popularity scoring requires participants to form usable expectations about what others will say. It is validated on single questions inside one frame. A grant panel spanning methodologists, practitioners and affected communities does not share a prior. How badly the method degrades there is unmeasured, and measuring it is the most valuable open problem in this design.

Slow flaws go unpaid. The panel pays for what can be assessed now. A defect that becomes visible in five years earns nothing when raised. Delayed re-scoring would fix it and would reintroduce an incentive to stay quiet.

Reviewer supply. Everything here assumes enough competent people keep turning up to read submissions properly. Payment and public standing are the answer offered; whether they are sufficient is untested, and this is what usually kills projects of this shape.

Coordinated attack. Open entry plus payment invites rings of accounts attacking each other for score. Random panel draws, paired reports and multi-task scoring make this expensive, slow and visible after the fact. They do not make it impossible, and this paper will not claim otherwise.

7. Retraction conditions

Stated in advance. If any is met, the claim above it is withdrawn in public rather than defended.

  1. If surprising-popularity scoring fails to recover informed-minority judgments on panels whose members do not share a background, and no repair is offered, the scoring rule in §3 is withdrawn and the mechanism with it.

  2. If critiques rated low on substance by independent expert assessment are paid systematically more than those rated high, the claim that the mechanism will not pay for theatre is false and the rule is retired rather than tuned.

  3. If contributions that confirm a proposal earn materially less than contributions that refute one, after controlling for panel expectation, the symmetry claim in §3 is withdrawn.

  4. If realised funder exposure in any round exceeds b · ln(n), the arithmetic in §3 is wrong.

  5. If the first comparison against randomised outcomes shows panel scores and market prices performing no better than an unweighted baseline, §5 is withdrawn as a contribution and published as a negative result with the data attached.

The fifth is the one that matters, and the one this paper would least like to lose.

8. What this does not claim

No component here is new. Surprising popularity, bounded-loss market scoring and partial lotteries are all someone else’s, cited below and working. The contribution is the composition and the observation in §5.

No incentive-compatibility result is claimed for the composed system. Each layer’s guarantees come from its source; how they behave together under strategic play across layers is unproven, and that is the first thing a referee should attack.

Nothing here has been run. This is a specification published so it can be broken, which is the only way to publish it that is consistent with what it argues.


References

Herbert, D. L., Barnett, A. G., Clarke, P., & Graves, N. (2013). On the time spent preparing grant proposals: an observational study of Australian researchers. BMJ Open, 3(5).

Hanson, R. (2003). Combinatorial information market design. Information Systems Frontiers, 5(1), 107–119.

Liu, M., et al. (2020). The acceptability of using a lottery to allocate research funding: a survey of applicants. Research Integrity and Peer Review, 5(3).

Othman, A., & Sandholm, T. (2010). Decision rules and decision markets. Proceedings of AAMAS 2010.

Pier, E. L., Brauer, M., Filut, A., Kaatz, A., Raclaw, J., Nathan, M. J., Ford, C. E., & Carnes, M. (2018). Low agreement among reviewers evaluating the same NIH grant applications. Proceedings of the National Academy of Sciences, 115(12), 2952–2957.

Prelec, D. (2004). A Bayesian truth serum for subjective data. Science, 306(5695), 462–466.

Prelec, D., Seung, H. S., & McCoy, J. (2017). A solution to the single-question crowd wisdom problem. Nature, 541(7638), 532–535.

Shnayder, V., Agarwal, A., Frongillo, R., & Parkes, D. C. (2016). Informed truthfulness in multi-task peer prediction. Proceedings of the 2016 ACM Conference on Economics and Computation.

Witkowski, J., & Parkes, D. C. (2012). A robust Bayesian truth serum for small populations. Proceedings of the 26th AAAI Conference on Artificial Intelligence.


Competing interests

The author is the principal of Aux Labs LLC and of The Switchboard, its nonprofit research arm. The Switchboard would hold the rules of any instance of this mechanism; Aux Labs intends to enter such instances as a competing submitter, with entries labelled and rule-change authority forfeited at launch. This is a governance weakness of the proposal and is disclosed here rather than in a footnote.