NextConsensus
Discuss a Program

Performance

How forecasts are evaluated

This page will eventually hold the numbers. It is written now, before there are any, so the standard is fixed while nothing is riding on it — what gets measured, in what order, and what may never be shown on its own.

Method v0.1.0 Status Retrospective
Specification

There are no results here yet, because nothing has resolved. What follows is the measuring stick, published in advance. Numbers appear only after independent evaluation against a defined cohort — never before.

Primary

Proper scoring rule

Brier score or log loss, always against named baselines. This is the headline number, and it only gets published once it has been earned.

Metric Status
Brier score not-yet
Log loss not-yet
Secondary

Calibration

Whether the numbers mean what they say. If a system is calibrated, the things it calls 70% happen about seventy times in a hundred — no more, and no less.

Metric Status
Reliability curve not-yet
Expected calibration error not-yet
Calibration slope and intercept not-yet
Operational

Threshold utility

Whether a forecast is any use at the point you would actually act on it. Lead time without precision is just a machine for generating false alarms.

Metric Status
Precision at threshold not-yet
Recall at threshold not-yet
Alert burden not-yet
Useful lead time not-yet
Diagnostic

Segmented performance

The same results split by authority, decision type, horizon, therapeutic area, and whether the forecast was made in advance or reconstructed from history. Averages hide where a method is weak.

Metric Status
Performance by authority not-yet
Performance by event class not-yet
Performance by horizon not-yet
Performance by therapeutic area not-yet
Prospective vs. retrospective breakdown not-yet

Every published metric must state

Every number published on this site will carry these fields.

What this page will never do

Report lead time without precision

A system can manufacture long lead time by repeatedly estimating transitions. Lead time is only meaningful alongside precision and alert burden.

Show calibration without cohort definition

Calibration is cohort-specific. A system can appear well-calibrated on one event class while failing on another. Pooling event classes is a category error.

Display retrospective results as live performance

Retrospective results are labeled retrospective. Prospective results are labeled prospective. Conflating them is a credibility violation.

Claim improvement without baselines

A Brier score of 0.15 is meaningless without knowing what the baseline achieves. Every metric is compared against named baselines.