Evaluating a ball carrier in football is a counterfactual problem. A running back gains eight yards. Was that good? It depends on what would have happened otherwise — what a typical player would have gained from the same position, against the same defense, at the same moment. The comparison isn't to an average. It's to a distribution of alternatives.
Bayesian multilevel step-and-turn models generate that distribution. At each frame of tracking data, the model decomposes player movement into two components: step length (how far) and turn angle (which direction). Both are modeled hierarchically — player-specific tendencies nested within play-level contexts. The posterior predictive simulation then generates thousands of hypothetical ball carrier trajectories from the same starting conditions, each representing what a plausible alternative player might have done.
The evaluation becomes geometric. The observed trajectory is one path through a cloud of simulated alternatives. Performance metrics emerge from the comparison: how much more (or less) yardage did the real player gain relative to the simulated population? How often did the real player's path fall in the tail of the distribution?
The structural contribution is the decomposition itself. Step-and-turn models come from animal movement ecology, where they describe foraging paths of caribou and albatrosses. Applying them to NFL tracking data doesn't just borrow a technique — it reframes what player evaluation measures. The player isn't being compared to a single replacement. They're being compared to a generative process that captures the full space of what running looks like, and the evaluation falls out as a tail probability rather than a point estimate.