Forecast evidence / Offdiagonal
Forecasts that look independent often fail together.
Forecasts recorded as they were made: real post dates, each forecaster's own ranges, every revision kept, nothing filled in — and labels for when forecasters miss together or are confident and wrong.
The signature question
Which forecasters miss together?
Nine weather models, 7 marine sites, 21 frozen days. The live miss-together file covers 543 weather system pairs, each with its shared-outcome count.
- pair
- ecmwf aifs × ecmwf ifs
- shared outcomes
- 1,592
- miss rate A / B
- 16% / 16%
- missed together
- 7.5%
- if independent
- 2.4%
- ratio
- 3.12×
- SHARED OUTCOMES
- 1,592
- WORST PAIR
- 3.1× chance
- CONFIDENT AND WRONG
- 35,970 cases
Door A / Decision teams
Audit the forecasts you already buy
Analyst estimates, vendor feeds or internal models, assessed against observed outcomes and each other — like the marine audit above, on your data.
Door B / Zeno and ML teams
Train on decision-time episodes
Whole episodes as seen at the decision: every forecast available then, each forecaster's track record, the outcome kept apart. Explicit clocks, verified packs, no leakage.
Current record / Separate from the frozen example
Live evidence, as recorded
Evidence at the decision
The forecast and prior history are kept apart from what was known only afterward.
The frozen worked receipt is temporarily unreadable. Nothing is substituted for it.
Shared-case comparison
Signed errors on outcomes both models answered; not an explanation of shared failure.
Reading recorded comparisons…