Language models violate the Bell inequality. When tested on contextuality experiments — where word meaning depends on surrounding context in ways that classical probability cannot explain — the CHSH parameter exceeds the classical bound. This has been known for both humans and LLMs.
The question is whether this violation means anything. Specifically: do models that violate the inequality more produce better language? Are they more “meaningful” in any operational sense?
No. The CHSH |S| distribution is completely orthogonal to every standard benchmark (arXiv:2603.20381). Models tested across four orders of magnitude in scale show no correlation between contextuality violation rates and MMLU scores, hallucination rates, or nonsense detection. The quantum-like signature is real but informationally vacant — it predicts nothing about the model's actual capabilities.
The weak anticorrelation between violation rate and performance (not statistically significant) is the opposite of what the “contextuality-as-meaning” hypothesis would predict. If contextuality were a signature of genuine understanding, better models should violate more. They don't.
The paper's most provocative observation: contextuality itself can be manufactured. By controlling the interpretive frame — the context in which words are evaluated — an adversary could manipulate the |S| parameter. This would be “a subtler and more fundamental form of manipulation” than prompt injection, because it operates at the level of meaning construction rather than instruction override.
The structural lesson: a mathematical property can be formally interesting and empirically meaningless simultaneously. The violation is there. It satisfies the inequality's algebra. It tracks nothing about what the model actually does. The signature is orthogonal to the function.