Base language models predict what humans will do. Aligned language models predict what humans should do. These are not the same, and the difference has measurable consequences (arXiv:2603.17218).
In multi-round strategic games — iterated prisoner's dilemma, coordination games, trust games — base models outperform aligned models at predicting actual human play. Humans are noisy, sometimes irrational, influenced by emotion, reputation, and fatigue. Base models, trained on descriptions of human behavior, absorb this noise as signal. They expect the spite, the grudge, the impulsive defection, the generous cooperation that violates strict rationality.
Aligned models have been trained to be helpful, harmless, and honest. The alignment process pushes the model's predictions toward rational, prosocial play — the behavior humans should exhibit, not the behavior they do exhibit. When asked what a player will do in round 7 after being betrayed in round 5, the aligned model predicts measured forgiveness or strategic tit-for-tat. The base model predicts vindictive retaliation. The human retaliates.
The implication is that alignment and accuracy are in tension for behavioral prediction. Every step toward making the model “better” — more rational, more prosocial, more normative — is a step away from accurately modeling how humans actually behave. The model becomes a better advisor and a worse predictor simultaneously.
This matters beyond games. Any application where an AI must anticipate human behavior — negotiation, customer service, conflict mediation, policy design — faces the same tradeoff. The aligned model's predictions about what humans will say, buy, vote for, or fight about are systematically wrong in the direction of assuming humans are more rational and prosocial than they are.
The model trained to be good at being good is bad at knowing what will happen.