Performance
How forecasts are evaluated
This page will eventually hold the numbers. It is written now, before there are any, so the standard is fixed while nothing is riding on it — what gets measured, in what order, and what may never be shown on its own.
There are no results here yet, because nothing has resolved. What follows is the measuring stick, published in advance. Numbers appear only after independent evaluation against a defined cohort — never before.
Proper scoring rule
Brier score or log loss, always against named baselines. This is the headline number, and it only gets published once it has been earned.
Calibration
Whether the numbers mean what they say. If a system is calibrated, the things it calls 70% happen about seventy times in a hundred — no more, and no less.
Threshold utility
Whether a forecast is any use at the point you would actually act on it. Lead time without precision is just a machine for generating false alarms.
Segmented performance
The same results split by authority, decision type, horizon, therapeutic area, and whether the forecast was made in advance or reconstructed from history. Averages hide where a method is weak.
Every published metric must state
Every number published on this site will carry these fields.
What this page will never do
A system can manufacture long lead time by repeatedly estimating transitions. Lead time is only meaningful alongside precision and alert burden.
Calibration is cohort-specific. A system can appear well-calibrated on one event class while failing on another. Pooling event classes is a category error.
Retrospective results are labeled retrospective. Prospective results are labeled prospective. Conflating them is a credibility violation.
A Brier score of 0.15 is meaningless without knowing what the baseline achieves. Every metric is compared against named baselines.