The Fractions Skill Score is one of the most widely used metrics for evaluating spatial weather forecasts. It works by comparing forecast and observed fields within expanding neighborhoods — if the forecast gets the right answer within a neighborhood of size r, it scores well. The intuition is that a forecast that's spatially displaced but otherwise correct should score better than one that misses entirely.
Antonio demonstrates that the FSS systematically rewards smoother forecasts. A blurrier forecast — one that spreads its probability over a larger area — scores higher than a sharp forecast with the same spatial displacement error, because the blurring increases the overlap between forecast and observation neighborhoods. The metric has a built-in preference for hedging.
The bias is structural, not accidental. The FSS also tends to assign higher scores to forecasts that over-predict mean frequency — forecasts that say “rain” too often score better than forecasts that say “rain” too rarely, holding other errors equal. The theoretical analysis confirms what practitioners have suspected: the FSS should be used with percentile thresholds (which normalize the frequency) rather than raw thresholds (which don't).
The double penalty problem — where a displaced forecast is penalized both for the miss at the observed location and the false alarm at the forecast location — affects the related Brier Divergence Skill Score more severely than the FSS. The FSS is more robust to this issue, which partly explains its popularity.
A standard metric, used across national weather services worldwide, theoretically confirmed to reward the wrong thing. The fix is known (percentile thresholds). The question is whether the metric's users know they need it.