Forecast evidence / Offdiagonal

Forecasts that look independent often fail together.

Forecasts recorded as they were made: real post dates, each forecaster's own ranges, every revision kept, nothing filled in — and labels for when forecasters miss together or are confident and wrong.

The signature question

Which forecasters miss together?

Nine weather models, 7 marine sites, 21 frozen days. The live miss-together file covers 543 weather system pairs, each with its shared-outcome count.

Inspect full report
cma grapesecmwf aifsecmwf ifsgemgfsiconjmameteofranceukmocma grapesecmwf aifsecmwf ifsgemgfsiconjmameteofranceukmo
pair
ecmwf aifs × ecmwf ifs
shared outcomes
1,592
miss rate A / B
16% / 16%
missed together
7.5%
if independent
2.4%
ratio
3.12×
Each cell is one pair of models. Darker means they missed the same outcome more often than chance alone would predict. Select a cell for the exact counts. On small screens, drag sideways to see every column.
SHARED OUTCOMES
1,592
WORST PAIR
3.1× chance
CONFIDENT AND WRONG
35,970 cases

Door A / Decision teams

Audit the forecasts you already buy

Analyst estimates, vendor feeds or internal models, assessed against observed outcomes and each other — like the marine audit above, on your data.

Door B / Zeno and ML teams

Train on decision-time episodes

Whole episodes as seen at the decision: every forecast available then, each forecaster's track record, the outcome kept apart. Explicit clocks, verified packs, no leakage.

Current record / Separate from the frozen example

Live evidence, as recorded

Evidence at the decision

The forecast and prior history are kept apart from what was known only afterward.

The frozen worked receipt is temporarily unreadable. Nothing is substituted for it.

Source provenance

Shared-case comparison

Signed errors on outcomes both models answered; not an explanation of shared failure.

Reading recorded comparisons…

Explore coverage