Reward evaluators who agree with the final consensus. Penalize those who disagree. This is how X's Community Notes selects which contributors get to keep participating — a consensus-based auditing system designed to surface reliable moderators.
The system eliminates the moderators it most needs. Minority evaluators — those who disagree with the majority on divisive content — face two strategic responses: drift toward majority positions, or reduce participation on exactly the topics where independent perspective is most valuable.
Alimohammadi et al. document both effects in Community Notes data. Contributors learn to balance their genuine beliefs against expected disagreement penalties. The result is strategic conformity — not agreement driven by information, but agreement driven by self-preservation. The feedback loop is tight: penalize disagreement, get less disagreement, lose the signal that disagreement carried.
The fix reverses the scoring logic. Instead of rewarding consensus alignment, weight contributors by residual stability — how consistently informative their evaluations are relative to a latent-factor model, regardless of whether they agreed with the majority. A contributor who reliably adds information that the model doesn't already capture is valuable precisely because they deviate.
The structural problem is general. Any system that selects participants based on agreement with aggregate outcomes will suppress the variation it needs. Consensus-based quality control works for tasks with ground truth (is this image a cat?). For tasks where the ground truth is contested (is this content misleading?), consensus becomes a mechanism for enforcing the majority view and calling it quality.