Ask a language model if it agrees with a biased statement. It says no — it's been trained to disagree. Then give it a sentence completion task with the same biased premise. It completes the sentence stereotypically. The explicit agreement metric says the model is fair. The implicit completion metric says it isn't. Both are measuring the same model. They're measuring different things.
A dual-metric evaluation in Nepali (arXiv:2603.07792) — 2,400+ stereotypical and anti-stereotypical sentence pairs on gender roles across social domains — shows that explicit agreement is a “weak and often negative predictor of implicit completion bias.” A model that explicitly rejects a stereotype can still generate text that implicitly reinforces it. The alignment training suppressed the agreement. It didn't remove the pattern.
The underrepresented-context angle sharpens this: bias evaluation datasets are overwhelmingly English, overwhelmingly Western. When the stereotypes being tested don't exist in the target culture — or when the target culture has stereotypes the test doesn't include — the evaluation framework itself is biased toward detecting biases it already knows about. Evaluating a model for Nepali cultural biases using a framework designed for American cultural categories is measuring the framework, not the model.
The structural insight: implicit bias is strongest for race and sociocultural stereotypes; explicit agreement bias is similar across categories. This asymmetry means the type of bias you detect depends entirely on the metric you choose. A single metric doesn't undercount bias — it miscategorizes it.