The NRCS runoff forecast report card
The statewide verdict
Across all 69 graded points, the median point’s April even-odds forecast (the “50% exceedanceA forecast stated as odds of AT LEAST this much: the 90% exceedance value is a near-floor the river beats nine years in ten, the 10% value a near-ceiling it reaches one year in ten, and 50% is even odds -- as likely to come in above as below.” number — as likely to come in above as below) misses the delivered volume by 24% in a typical year. The stated likely range (the 90% value the river should beat nine years in ten, up to the 10% value it should reach one year in ten) actually contains the outcome 75% of the time at the median point — a perfectly calibrated forecast would say 80%. 36 points run more than 5% high on average and 5 run more than 5% low. New to these numbers? How to read a runoff forecast →
This is a grade of the forecasts, never of the forecasters — these are hard rivers, and a 17–25% typical miss on an April guess at a summer’s runoff is real skill. What this page adds is the number NRCS does not publish beside each forecast: how it has verified, point by point, over five decades.
The dry-year problem
Split every point’s record into its own driest, middle and wettest thirds and the forecasts stop looking even-handed: in dry years the April even-odds forecast runs a median of +36% high; in wet years it runs 10% low. The stated 10–90% band, nominally right 80% of the time, holds in only 67% of dry years at the median point. The operational reading is uncomfortable: the years a water manager most needs the number are the years it is most likely to promise water that never comes. Regression to the mean and thirsty post-drought soils are the textbook explanations — the size of the effect here is the measurement.
Warm bars are statewide-dry years (median point delivered under 80% of its normal year), blue bars wet ones. The driest years on record here — 1977, 2002, 2012, 2018, 2021 — all sit far above the line; 2002’s statewide median miss was +98%. And the skill is not trending better: 1950s 25%, 1960s 17%, 1970s 19%, 1980s 23%, 1990s 23%, 2000s 24%, 2010s 23%, 2020s 29% median miss — the 2020s are so far the roughest decade in the archive, which is what a snow–runoff relationship degrading under aridification would look like.
Where the drought bias lives
Each forecast point, colored by its dry-year bias -- how far the April forecast typically runs above (red) or below (blue) what the river delivers in that point’s own driest third of years. Larger dots carry more graded years; hollow gray dots are the flagged ⚠ points whose numbers likely grade a naturalized-vs-gaged mismatch rather than the forecast. Click any dot for its numbers.
What a month of waiting buys
The same points, graded at every issue date: the January outlook misses by a median 29%, June’s by 20%. Most of the improvement arrives late, when the snowpack’s fate is already sealed — which is exactly why the dry-year problem above matters: the early-season number is the one decisions get made on, and it is the least certain.
Do the stated odds mean what they say?
Each forecast publishes five exceedance levels — "90% chance the volume is at least X". Pooling 1,937 point-years that published all five: a value stated with 90% confidence was actually reached 75% of the time; one stated at 10% was reached 9%. A calibrated forecast plots on the diagonal.
Accuracy against calibration
Right is a bigger typical miss; the dashed line is where a calibrated band should sit. Larger dots carry more graded years. Every dot is named on hover and listed in the table below. Points whose typical miss exceeds 100% are left off this chart (and flagged ⚠ below) — a miss that size usually means the forecast and the gage measure different water (see the method notes).
All 69 points, best typical miss first download CSV ↓
Method, sources and honest limits
Forecasts: NRCS National Water and Climate Center archive (AWDB), element SRVO — the seasonal streamflow-volume forecast, exceedance values at 90/70/50/30/10%. For each point and year, the LAST issue published in April is graded, against the period that issue itself declared (NRCS has changed the window over the decades; each year verifies against its own). Observed: the daily record DWR publishes for the forecast point’s own gage, summed over the declared period; years with under 90% daily coverage are dropped, never silently summed. Median miss is the median over years of |50% forecast − observed| as a percent of observed. Bias is the median signed error: positive means the forecast typically ran high. In-band is the share of years the observed volume fell between the 90% and 10% exceedance values — nominally 80% — computed only over years that HAVE a published band (early decades often carried only the 50% number). Dry-year bias is the median signed error over the driest third of that point’s own gradable years (points with fewer than nine show a dash). The yearly chart shows only years with ten or more graded points; its dryness index is the median across points of the year’s volume over that point’s own median year. Only points whose gage record this site holds are graded (69 of ~80 Colorado forecast points), and points with fewer than five gradable years are left out entirely. The one caveat that matters: at some points NRCS forecasts an ADJUSTED (naturalized) volume -- corrected for upstream storage and diversions -- while the gage records the regulated flow that actually passed. Where the two diverge badly, the "miss" is a definitional mismatch, not a forecast failure; points whose typical miss exceeds 100% almost certainly fall in this class and are flagged ⚠ rather than removed, because deciding which points NRCS naturalizes is itself research this page has not yet done. Skill archive refreshed monthly; this rendering 2026-08-31. Per-point year-by-year charts live on each district’s Snow & runoff tab.