Better calibrated probabilities
A forecast of 70% should happen roughly seven times in ten. Calibration is measured explicitly: expected calibration error fell from 7.51% to 6.8%.
White Christmas USA combines official NOAA snow-depth observations with USGS terrain and geographic data. The current model was trained and calibrated through 2020, then assessed against 12,588 observations from 2021–2025 that were excluded from development.
+3.18%
Lower probability error
Brier skill versus NOAA interpolation on the independent 2021–2025 test set.
12,588
Independent outcomes
Station-by-Christmas results from 2021–2025, excluded from training and tuning.
0.870
ROC–AUC discrimination
About 87 in 100 correct rankings of a white versus non-white Christmas observation.
9/9
Geographic tests improved
Every Census-region and spatial-block holdout improved on the benchmark.
Independent testing
Using the same records to build and assess a model can produce misleading results. We divided the observations by time so that the final assessment used five Christmases the model had not encountered during training or calibration.
1991–2015
Estimate relationships between snow history, terrain and geography.
2016–2020
Choose model settings and adjust the probability scale.
2021–2025
Compare the completed model with the NOAA baseline on held-out observations.
Result: lower Brier error, lower log loss, stronger ranking and better calibration than the NOAA four-station benchmark.
A forecast of 70% should happen roughly seven times in ten. Calibration is measured explicitly: expected calibration error fell from 7.51% to 6.8%.
Every prediction draws on nearby stations, elevation, local relief, latitude, longitude, coast distance and Great Lakes proximity. The target station is excluded from its own predictors during validation.
The model was separately tested across all four Census regions and five spatial blocks. It improved on the benchmark in all nine tests; the smallest improvement was 0.08%.
Method
The system is sophisticated underneath, but its job is simple: estimate the chance of at least one inch of snow on the ground at 07:00 local time on Christmas morning.
Stage 01
Official NOAA 1991–2020 Climate Normals provide the foundation. GHCN-D station observations add year-by-year ground truth, with traceable source stations and weights behind every result.
Stage 02
USGS 3DEP elevation and local relief help distinguish a mountain town from a nearby valley. Coastline and Great Lakes distance capture geographic snow regimes that a simple nearest-station average misses.
Stage 03
A monotonic histogram gradient-boosting model combines these inputs. It uses an ensemble of constrained decision trees to identify interactions while preserving sensible relationships between the predictors and snow probability.
Stage 04
A new version must improve probability accuracy, preserve calibration and avoid geographic failure. A model using 21 years of SNODAS analysis was not adopted because its test error was 0.41% worse than the NOAA baseline.
Inside the forecast window
Once Christmas falls within the medium-range forecast window, the historical estimate is blended with NOAA's global ensemble forecast. The live weights were selected using a separate five-Christmas reforecast test and remain deliberately conservative.
+36.7%
day-one retrospective skill
Versus climatology across 13,891 held-out station outcomes.
+14.6%
or better at every lead
Positive skill at days 1, 3, 5, 7 and 9 on the held-back reforecast years.
Important context: the retrospective GEFS predictor is snow-water equivalent, while the operational feed provides direct snow depth. The reforecast result informs conservative blend weights; it is not presented as direct operational validation.
65,664
local outlooks
32,024
cities, towns & communities
33,640
ZIP Code Tabulation Areas
4,551
NOAA climate stations
Limits and interpretation
Each result is a statistical estimate for a representative point. ZIP results use a point within the Census ZIP area, and mountain microclimates can vary over very short distances.
The model is demonstrably stronger than the NOAA interpolation benchmark used here. It has not been tested against every private forecast system, so we do not claim it is universally the most accurate model in America.
Source data and model versions remain traceable. Models that fail the validation criteria are not released, and each Christmas adds more direct snow-depth evidence for the operational forecast.
Inspect the full model statisticsExplore the national map or search down to a city, town, village or ZIP area.
Questions about the methodology? Get in touch.