Gestalt grouping principles — proximity, similarity, common fate, good continuation — describe how the human visual system organizes the world into coherent objects. These principles were articulated a century ago and have been confirmed across cultures and species. They appear to be fundamental properties of visual cognition.
Self-supervised vision transformers, trained only to predict masked patches or match augmented views, spontaneously develop the same grouping strategies. No labels. No explicit grouping objective. No human feedback on what constitutes a “group.”
The match is quantitative: when evaluated on psychophysical tasks designed to measure perceptual grouping in humans, the transformer representations produce grouping patterns that closely track human performance across multiple Gestalt principles. The model doesn't merely cluster features — it clusters them in the same way that human observers do, prioritizing the same cues in the same relative ordering.
The result suggests that these organizational principles are not arbitrary cognitive heuristics but statistical regularities of natural images that any sufficiently capable visual system will discover. The transformer finds them because they're there — because proximity, similarity, and good continuation are genuine properties of how visual scenes are structured, and the self-supervised objective implicitly rewards learning them.
The flip side is that every principle the model fails to learn points to something the visual statistics alone can't explain — a contribution of the visual system's architecture or developmental history that goes beyond the data. If proximity and similarity emerge from statistics but common fate does not (or does so weakly), that gap tells you where biology adds something that pure learning from images does not provide.
The principles aren't instructions. They're what you converge on when you look at the world hard enough.