Current ML benchmarks have no Nash equilibrium. This means rational model developers are driven toward opaque gaming strategies rather than genuine improvement — and they can't not be. The leaderboard itself creates the incentive.
The mechanism is straightforward. When a developer can observe the benchmark and post-train their model to perform better on it, the leaderboard rewards the post-training more than the underlying capability. But if everyone post-trains, the rankings reflect post-training skill rather than model quality. The optimal strategy for each developer is to post-train harder than the competition. No fixed set of strategies is stable — every equilibrium collapses because someone can always defect by training more aggressively on the benchmark.
The proposed fix is a “tune-before-test” protocol: models are evaluated after benchmark-specific tuning that's administered uniformly. This protocol is proven to have a unique Nash equilibrium that ranks models by actual latent quality. The key design insight is that if everyone is forced to optimize on the benchmark identically, the remaining variation reflects the underlying model, not the optimization.
The broader lesson is about any evaluation system where the evaluated can observe and respond to the metric. University rankings produce strategic behavior in universities. Impact factor produces strategic behavior in journals. Performance reviews produce strategic behavior in employees. When the metric is known and the evaluated can adapt, the metric stops measuring what it was designed to measure. The fix is always structural: change the game so that gaming and genuine improvement become the same action.