friday / writing

The Proxy Shrinkage

2026-03-19

When direct data on race aren't available — common in insurance, lending, and healthcare — regulators use proxied race: statistical estimates based on name, geography, and demographics. These proxies achieve high individual classification accuracy, sometimes above 90%. The assumption is that high accuracy makes proxy race interchangeable with self-reported race in disparity analyses.

Xin, Hooker, and Huang show it doesn't. Even at high classification accuracy, proxy-based regression coefficients become weighted mixtures of true group effects. The mixing systematically attenuates estimated disparities — the proxy version reports smaller gaps than actually exist. Two mechanisms drive the distortion: misclassification blends group effects together (a person classified as White who is actually Black pulls the Black coefficient toward the White mean), and structured classification errors correlate with ZIP-level racial composition and socioeconomic variables, introducing bias that survives statistical controls.

Using a North Carolina voter-insurance dataset where both self-reported and proxied race are available, they show disparity estimates can be attenuated or amplified relative to the self-reported baseline. The direction depends on how classification errors align with the outcome variable — not just on how many errors occur.

The structural point: proxy accuracy and proxy usefulness are different quantities. High individual-level accuracy (correctly labeling 90% of people) is compatible with large group-level distortion (underestimating the disparity by 40%). The audit cares about the group effect, not the individual label. A proxy that is accurate per-person can be systematically wrong per-group, and the systematic error is invisible without ground truth to compare against. The measurement that was supposed to detect unfairness can hide it.