Example report · frozen
Marine forecast audit: nine models
Nine models forecast the same 1,592 marine outcomes. Compare their ranges and the mistakes they share.
2026-08-28 to 2026-09-27 · 21 days · 7 sites · minimum 200 cases per comparison
Why this shared cohort?
Nine global numerical weather models forecasting 2 m temperature (and three other variables) at 7 marine sites. It is the largest fully shared cohort in our frozen, checksummed evidence: every model forecast the same 1,592 outcomes, each forecast was available before its outcome, and nothing is excluded for missing availability. That makes pairwise comparisons fair.
Finding 1
Every model's 90% range missed more often than 10% — from 16% to 33% of the same 1,592 outcomes.
What it could change: Ranges this wide of their nominal coverage warrant recalibration before being used for thresholds or alerts.
Finding 2
Some pairs miss together far more than chance alone would give — up to 3.1× the independence reference.
- pair
- ecmwf aifs × ecmwf ifs
- shared outcomes
- 1,592
- miss rate A / B
- 16% / 16%
- missed together
- 7.5%
- if independent
- 2.4%
- ratio
- 3.12×
What it could change: Pairs that miss together add less protection than their count suggests. Test whether a provider with different errors improves a combined forecast before relying on agreement.
Interpretation and limits
Joint misses partly reflect hard cases and inaccurate models. The independence rate is a descriptive reference only; this does not show cause, nor that any provider is redundant.
Finding 3
The problem is broad, not local: the median model misses 22–27% in every variable and lead time. The gap between models is wider than the gap between variables.
| variable | lead 0-24h | lead 25-72h |
|---|---|---|
| temperature 2m | median 28% · range 16%–33% · n 1,592 | median 27% · range 20%–36% · n 1,045 |
| wind speed 10m | median 27% · range 13%–32% · n 1,591 | median 27% · range 13%–33% · n 1,045 |
| surface pressure | median 25% · range 9%–40% · n 1,591 | median 27% · range 20%–44% · n 1,045 |
| precipitation | median 22% · range 16%–25% · n 1,591 | median 24% · range 22%–30% · n 1,045 |
What it could change: Choosing or reweighting models is likely to matter more than fixing one variable. Cells with too few cases would be hidden rather than shown as zero.
Not shown here
Available in a full audit
- Which forecasts beat a simple baseline? Available in audits; not shown here because this frozen extract stores misses, not point errors.
- Does adding another provider improve the combined forecast? Measured in audits on held-back matched outcomes with a combination rule fixed in advance. Not claimed from co-misses.
How strong is this
Limitations
- Ranges are constructed by Offdiagonal from the cross-model spread, not declared by the providers.
- 21 days at 7 sites: short history; results describe this cohort only.
- Outcomes near each other in time and place are not independent; case counts overstate effective sample size.
- A forecasting improvement does not by itself establish a financial or operational benefit.
Next step
Want this for your own forecasts?
The same structure applies to analyst estimates, vendor feeds or models. Tell us what you forecast and the decision it supports.
Request a forecast audit