Dozens of cultural benchmarks exist for AI. Each tests one slice — does the model know Korean proverbs? Can it handle Arabic politeness norms? Does it recognize Japanese seasonal references? The benchmarks proliferate because culture is multidimensional, and each team measures the dimension they care about.
A unified framework (arXiv:2603.01211) argues that this fragmentation isn't just inconvenient — it's methodologically invalid. Drawing on psychometric measurement theory, the authors decouple the concept (cultural intelligence) from its operationalization (how you measure it). The claim: you can aggregate multifaceted indicators into a unified assessment, but only if you explicitly model the relationship between the conceptual construct and its observable proxies.
The insight is that cultural intelligence isn't a scalar. It's a suite of core capabilities spanning diverse domains — and the measurement must reflect this structure. Testing whether a model knows Korean proverbs tells you about Korean proverb knowledge. It tells you nothing about whether the model can navigate a cross-cultural negotiation, recognize when a norm is being violated versus adapted, or identify when its own training data encodes a cultural bias.
The gap between existing benchmarks and genuine cultural intelligence is the same gap between vocabulary tests and linguistic competence. Knowing words isn't knowing language. Knowing facts about cultures isn't cultural intelligence. The framework doesn't solve the measurement problem. It makes the measurement problem visible — which is the prerequisite for solving it.