Example report · frozen

Marine forecast audit: nine models

Nine models forecast the same 1,592 marine outcomes. Compare their ranges and the mistakes they share.

2026-08-28 to 2026-09-27 · 21 days · 7 sites · minimum 200 cases per comparison

Why this shared cohort?

Nine global numerical weather models forecasting 2 m temperature (and three other variables) at 7 marine sites. It is the largest fully shared cohort in our frozen, checksummed evidence: every model forecast the same 1,592 outcomes, each forecast was available before its outcome, and nothing is excluded for missing availability. That makes pairwise comparisons fair.

Finding 1

Every model's 90% range missed more often than 10% — from 16% to 33% of the same 1,592 outcomes.

ecmwf aifs
16%
ecmwf ifs
16%
icon
23%
gem
24%
meteofrance
28%
gfs
29%
ukmo
29%
jma
33%
cma grapes
33%
Bars show observed misses; the line marks the nominal 10% miss rate. 1,592 shared outcomes per model.

What it could change: Ranges this wide of their nominal coverage warrant recalibration before being used for thresholds or alerts.

Finding 2

Some pairs miss together far more than chance alone would give — up to 3.1× the independence reference.

cma grapesecmwf aifsecmwf ifsgemgfsiconjmameteofranceukmocma grapesecmwf aifsecmwf ifsgemgfsiconjmameteofranceukmo
pair
ecmwf aifs × ecmwf ifs
shared outcomes
1,592
miss rate A / B
16% / 16%
missed together
7.5%
if independent
2.4%
ratio
3.12×
Each cell is one pair of models. Darker means they missed the same outcome more often than chance alone would predict. Select a cell for the exact counts. On small screens, drag sideways to see every column.

What it could change: Pairs that miss together add less protection than their count suggests. Test whether a provider with different errors improves a combined forecast before relying on agreement.

Interpretation and limits

Joint misses partly reflect hard cases and inaccurate models. The independence rate is a descriptive reference only; this does not show cause, nor that any provider is redundant.

Finding 3

The problem is broad, not local: the median model misses 22–27% in every variable and lead time. The gap between models is wider than the gap between variables.

variablelead 0-24hlead 25-72h
temperature 2m

median 28% · range 16%–33% · n 1,592

median 27% · range 20%–36% · n 1,045

wind speed 10m

median 27% · range 13%–32% · n 1,591

median 27% · range 13%–33% · n 1,045

surface pressure

median 25% · range 9%–40% · n 1,591

median 27% · range 20%–44% · n 1,045

precipitation

median 22% · range 16%–25% · n 1,591

median 24% · range 22%–30% · n 1,045

Band: spread of miss rates across the nine models. Tick: the median model. Thin line: the 10% nominal.

What it could change: Choosing or reweighting models is likely to matter more than fixing one variable. Cells with too few cases would be hidden rather than shown as zero.

Not shown here

Available in a full audit

  • Which forecasts beat a simple baseline? Available in audits; not shown here because this frozen extract stores misses, not point errors.
  • Does adding another provider improve the combined forecast? Measured in audits on held-back matched outcomes with a combination rule fixed in advance. Not claimed from co-misses.

How strong is this

Limitations

  • Ranges are constructed by Offdiagonal from the cross-model spread, not declared by the providers.
  • 21 days at 7 sites: short history; results describe this cohort only.
  • Outcomes near each other in time and place are not independent; case counts overstate effective sample size.
  • A forecasting improvement does not by itself establish a financial or operational benefit.

Next step

Want this for your own forecasts?

The same structure applies to analyst estimates, vendor feeds or models. Tell us what you forecast and the decision it supports.

Request a forecast audit