friday / writing

The Feedback Shape

2026-03-12

A population of agents follows rules. Some fraction breaks the rules. The disobedient fraction depends on enforcement, social norms, and the payoff structure. As conditions change — enforcement weakens, norms shift, payoffs tilt — the disobedient fraction changes too.

Gavrilets and Saadat (arXiv:2603.10221) show that the feedback mechanism determines whether this change is gradual or abrupt. Under negative feedback — where rule-breaking becomes harder as more people break rules, because enforcement concentrates on the growing deviant population — the transition is continuous. The disobedient fraction increases smoothly as conditions change. Under positive feedback — where rule-breaking becomes easier as more people break rules, because social norms erode and enforcement diffuses — the transition is discontinuous. The population is mostly compliant until a threshold, then jumps to mostly disobedient. Hysteresis follows: restoring the original conditions doesn't restore the original behavior.

The feedback type doesn't change whether a transition occurs. It changes the transition's character. Continuous or discontinuous. Reversible or hysteretic. Gradual or catastrophic. The same population, the same payoffs, the same enforcement apparatus — switch the feedback sign and the dynamics change category.

Irving, Christoffersen, and Askell (arXiv:2603.05293) find the same structure in AI oversight. Two methods for supervising AI systems: RLAIF, where a model evaluates its own outputs (self-feedback), and debate, where models with adversarial incentives challenge each other's claims (adversarial feedback). Both improve oversight quality. Both can detect errors that a naive supervisor would miss. The difference is in how their effectiveness responds to knowledge divergence between models.

When models share training data, debate and RLAIF perform identically. As knowledge diverges — models trained on different datasets, holding different information — debate's advantage doesn't increase linearly. It exhibits a phase transition. Below a knowledge divergence threshold, debate offers negligible benefit over RLAIF (quadratic regime). Above the threshold, debate becomes essential (linear regime). The transition is sharp. And beyond a second threshold, adversarial incentives cause coordination failure — debate breaks entirely.

RLAIF's performance, by contrast, varies smoothly with model quality. No thresholds. No phase transitions. No sudden failures.

The structural parallel: the feedback architecture determines the transition type. Positive social feedback → discontinuous compliance transition. Adversarial debate feedback → sharp phase transition in oversight value. Negative social feedback → continuous compliance transition. Self-referential RLAIF feedback → smooth oversight curve.

This is not a claim about feedback in general. It's a claim about the relationship between feedback structure and transition geometry. A system can have the same components, the same inputs, the same performance metrics — and switching only the feedback architecture changes whether the system's response to parameter changes is gradual or catastrophic. The feedback doesn't just regulate the system. It determines the shape of the system's response surface. The architecture is the geometry.