Real fish largely ignore virtual fish displayed on a screen. They swim away, lose interest, break formation. Yet a reinforcement learning algorithm trained on this uncooperative population still learned to guide fish schools toward target directions -- substantially outperforming both no-stimulus and simple heuristic baselines. The strategy that emerged was not brute persistence but adaptive timing: the virtual fish learned when to lead, when to pause, and when to reposition, calibrating its behavior to the partial and intermittent attention of the real animals.
The result inverts a common assumption about influence. Effective guidance does not require consistent followership. It requires a policy calibrated to inconsistent followership. The algorithm succeeded not because the fish reliably responded but because the algorithm learned the statistical structure of their unreliable responses -- which contexts produced brief compliance, which produced flight, which produced indifference -- and exploited the compliant windows. The fish were not trained; the influencer was trained to work with untrained subjects.
This reframes what it means to control a collective. The controller does not need to command; it needs to be opportunistic within the collective's own behavioral repertoire. The collective retains its autonomy at every moment, and the net directional effect emerges from the accumulation of small, well-timed nudges rather than from any single act of obedience. Influence, in this framework, is not the ability to be followed but the ability to exploit the moments when following happens to occur.
(arXiv:2603.16384)