offdiagonal
When a model predicted the world, what did it actually know — and was it right?
offdiagonal is a working laboratory that replays forecasts against reality from their decision cutoff. Nothing is interpolated, nothing is scored against information it could not have had.
receipts captured
0
outcomes resolved
0
producers tracked
0
sites observed
0
evidence cache temporarily unavailable — showing nothing rather than something stale
Error grows with how far ahead you ask
Mean absolute error against lead time, per producer and channel, computed from resolved receipts — not from any published accuracy claim.
lead-time curves unavailable — the summary cache has not been refreshed
What survives the replay rules
Most captured evidence does not qualify. Rows whose clock is assumed rather than measured are excluded and counted, never quietly kept.
clock basis of the admitted rows
derived by replay 1,316
No row yet carries a measured availability clock: today's admitted rows are eligible under replay assumptions, not proof of prospective availability. Measured clocks are being collected forward from now.
2 of 16 lineages contribute admitted rows across 12 calendar days · panel fp-panel-20260810-20260909 · training_eligible=false
How it works
- 01
Capture
Every prediction is recorded with its clock — what was emitted, when, and what was knowable at that moment.
- 02
Wait
Reality arrives. Outcomes are captured as separate, versioned evidence — never back-written into the forecast.
- 03
Replay
Each forecast is scored only against information available at its decision cutoff. Assumed or unknown clocks are excluded, not guessed.
- 04
Preserve
Receipts are immutable and append-only. Any result can be replayed, audited, and challenged.
Why it matters
Forecasts are everywhere; accountability is not. Models are compared on marketing claims and retrospective scores computed against information nobody had at the time. offdiagonal makes model behaviour measurable, comparable, and replayable — evidence infrastructure for anyone building, buying, or auditing predictive systems.
Said honestly: usable evidence is narrow today — the figures above show exactly how narrow. That narrowness is published, not hidden; the credibility is the point.
The full corpus is public and machine-readable.