AI agents in repeated economic interactions often fail to converge to Nash equilibrium. The standard fix is post-training alignment — retrain the agents to play strategically. But alignment requires controlling all the agents, which is impractical when they come from different developers with different training procedures.
This paper proves that alignment is unnecessary for a specific class of agents. “Reasonably reasoning” agents — those capable of forming beliefs about others' strategies from observation and learning to best respond to those beliefs — naturally converge to Nash-like play along almost every realized path, zero-shot, without any post-training. The convergence holds even when payoffs are unknown and agents observe only their own private stochastic rewards.
The result says that the capability to reason about others and adapt is sufficient for strategic stability. You don't need to engineer equilibrium behavior; it emerges from the reasoning process itself. The empirical validation covers repeated prisoner's dilemma through marketing promotion games. The implication for AI deployment: if agents can reason and observe, strategic equilibrium is an intrinsic property of interaction, not an extrinsic constraint to be imposed.