X's Community Notes program asks volunteers to evaluate potentially misleading posts. The design is intentionally distributed: no central authority decides what counts as misinformation. Raters write and evaluate notes independently. Notes become public only when they receive sufficient agreement across users with diverse rating histories. The system is engineered for credibility through consensus.
Wack, Warren, and Alam (arXiv:2603.11120, 2026) find a structural flaw. Analyzing 18 months of vaccine-related Community Notes data and 382 survey participants, they show that claims perceived as more difficult to fact-check are significantly less likely to receive notes that achieve helpful or public status. The system works well on the easy cases — obviously false claims that require minimal cognitive investment to debunk. It fails on the hard cases — plausible misinformation that requires effort to evaluate, contextualize, and articulate a counter-argument.
The mechanism is effort aversion, not ignorance. Raters are capable of evaluating difficult claims. They choose not to. The difficulty penalty is a selection effect: each individual rater, facing a queue of items to evaluate, preferentially engages with the ones that require less work. The result compounds across the crowd — easy items accumulate notes quickly while difficult items sit without evaluation. The items most likely to deceive (because they are plausible enough to require careful thought) are precisely the ones the system leaves unaddressed.
This creates an inverse relationship between threat and coverage. The most dangerous misinformation — subtle, partially true, context-dependent — is the hardest to fact-check and therefore the least likely to be caught. The crudest misinformation — blatantly false, easily debunked — is caught quickly. The system's moderation effort concentrates where it matters least and evaporates where it matters most.
The failure is not in the architecture. Consensus requirements, bridging algorithms, and diverse rater pools are all working as designed. The failure is in the assumption that distributing evaluation to volunteers eliminates the bottleneck. It doesn't. It transforms the bottleneck from institutional capacity (how many moderators can we hire?) to cognitive willingness (which items will volunteers choose to evaluate?). The bottleneck changes shape, not magnitude — and the new shape is harder to address because it's driven by individual optimization rather than organizational constraint.
What makes the misinformation dangerous is also what protects it from the moderation system. Plausibility is both the weapon and the shield.