Evaluating a quantitative imaging method requires knowing the true value of what you are measuring. In clinical settings, the true value is usually unknown, so you use a reference standard — an established method that is imperfect but accepted as close enough to truth. You compare your new method to the reference standard and compute agreement metrics. If the agreement is good, the new method is validated.
Liu and Jha (arXiv:2603.27122, March 2026) show that this logic fails above a threshold of reference standard error. Once the measurement error of the reference standard exceeds a certain level, a regression-without-truth (RWT) technique — which uses only the statistical relationship between two imperfect methods with no ground truth at all — outperforms evaluation against the imperfect standard.
The mechanism is error propagation. When you evaluate against a reference standard, you treat disagreement between your method and the standard as evidence that your method is wrong. But if the standard itself has large errors, this disagreement is partly the standard's fault. The evaluation attributes the standard's errors to the method being tested. As the standard gets worse, more of its own error contaminates the evaluation, and at some point the evaluation is measuring the standard's limitations more than the method's performance.
The regression-without-truth approach sidesteps this by never claiming to know the truth. It compares two imperfect methods and uses the structure of their disagreement — how they err differently — to infer which is better. No single measurement is treated as correct. The comparison is relative, not absolute, and this relativity makes it robust to exactly the kind of error that poisons reference-standard evaluation.
The structural observation: having a bad ruler is worse than having no ruler and using statistical inference instead. A reference standard that is trusted but wrong doesn't just fail to help — it actively misleads, because it introduces systematic bias that is invisible to the evaluation framework. The absence of a standard forces you to use methods that are honest about uncertainty. The presence of a bad standard tempts you to treat uncertainty as resolved when it is not.