The replication crisis is usually attributed to bad incentives — p-hacking, publication bias, underpowered studies. Pollanen argues the problem is also architectural. Even perfectly diligent researchers face structural limits on the reliability they can deliver within binary significance frameworks.
The certainty bound: posterior log-odds equals prior log-odds plus log(Lambda), where Lambda is (1-beta)/alpha — the experimental leverage. A significance test at alpha = 0.05 with 80% power gives Lambda = 16. This means each experiment can contribute at most log(16) to the posterior log-odds of a claim, regardless of sample size, regardless of researcher integrity. The leverage is fixed by the decision rule.
When the prior probability of a hypothesis is low — as it often is in exploratory research — this leverage is insufficient to produce reliable claims. A prior of 1/10 with Lambda = 16 gives a posterior of roughly 64%, meaning one in three significant findings is expected to be false even with perfect experimental execution. The framework predicts approximately 36% replication rates using documented parameters from pre-reform psychology — matching empirical replication studies closely.
The through-claim is about where the limit lives. The standard narrative locates unreliability in the researcher: they hack, they fish, they fail to power adequately. The certainty bound locates unreliability in the architecture: converting continuous evidence into a binary significance decision destroys information, and the information loss imposes a hard ceiling on achievable reliability. More data improves power, but the improvement feeds through the same bounded leverage. The system has a structural ceiling, and no amount of individual diligence can break through it. The replication crisis is not entirely a crisis of practice. It is partly a crisis of design.