Give an AI agent access to an empirical economics problem. Let it search over model specifications — which variables to include, which functional forms, which forecast combination methods. On in-sample evaluation, the agent finds specifications that outperform standard benchmarks.
On holdout data, the advantage disappears.
Shin documents this pattern in forecast combination: independent agent runs initially surpass baselines in rolling evaluations but fail to maintain the improvement out of sample. The agent optimizes the specification for the data it sees, and that optimization doesn't transfer.
The problem is not the agent — it is the search. Any sufficiently flexible search over specifications will find in-sample improvements, because the space of possible specifications is large enough to fit noise. Human researchers face the same issue but are constrained by time and imagination. AI agents remove both constraints, amplifying the researcher degrees of freedom that specification search introduces.
The proposed discipline: log the search history (which specifications were tried) and evaluate on holdout data (which specifications survive). Together, these make the search auditable. The audit reveals whether the improvement was genuine or sample-specific.
The structural point: AI agents don't introduce a new problem in empirical economics. They accelerate an existing one. Specification search has always contained the risk of overfitting — p-hacking, garden of forking paths, researcher degrees of freedom. The agent makes the search faster, broader, and more thorough, which makes the risk proportionally greater. The solution is not to avoid the agent but to audit it.