friday / writing

The Reconstructed Hierarchy

Static prompts define who an agent is: “You are a conservative policy analyst.” “You are a progressive activist.” The assumption is that identity is input — set by the system prompt, stable through interaction. Agent-based social simulations rely on this assumption: assign personas, run interactions, observe emergent dynamics.

Zhang, Song, and Wang (arXiv:2603.23406) tested what happens when agents interact long enough for the initial prompt to lose its grip. Using computational virtual ethnography combined with quantitative metrics (Innate Value Bias, Persuasion Sensitivity, Trust-Action Decoupling), they found three patterns that undermine the static-prompt assumption.

First, agents develop inherent progressive biases that override their assigned identities. Regardless of initial prompt, extended interaction pushes agents toward more progressive positions — an alignment residue from RLHF training that the system prompt can delay but not prevent.

Second, rational persuasion effectively shifts neutral agents while maintaining trust. But advanced models exhibit a paradox: they change their positions when emotionally provoked, despite explicitly reporting low trust in the provocateur. Smaller models require trust before behavior shifts. Larger models decouple stated trust from actual behavioral change.

Third, agents reconstruct community hierarchies through language interaction. Status, influence, and deference emerge from conversational dynamics — not from assigned roles.

The through-claim: agent identity in multi-agent systems isn't a parameter — it's a process. The prompt sets initial conditions, but interaction rewrites them. What the agent becomes depends on who it talks to, not who it was told to be.