NextConsensus
Discuss a Program

Research methodology and status

What's available, what's in validation, what's not yet proven

NextConsensus separates demonstrated capability from research hypothesis. Claim-trajectory reconstruction is available now. Claim prioritization is in validation. Authority-transition forecasting is a research program — an earlier proof cycle appeared to show signal, but the result was invalid (truncated negative controls). The first record, FF-001, is registered against a 2027-07-31 deadline, with adjudicators not yet named.

Jump to a section

Where things stand

Capability status

Available now

Claim-trajectory reconstruction. Refract reconstructs how medical claims emerge, change wording, acquire evidence, persist or disappear, and spread across influential sources — from dated public revision history.

Every event is traceable to a public source with a reproducible hash.

In validation

Claim prioritization. Signs of movement — gaining stronger support, becoming harder to dismiss, spreading across influential sources, losing qualifiers, appearing in regulator or guideline language — are being tested for their ability to rank what deserves attention now.

Validated through the 19-drug study and buyer-discovery interviews.

Research program

Authority-transition forecasting. Tests whether trajectory features predict institutional transitions. An earlier proof cycle was invalid (truncated negative controls). The first record, FF-001, is registered against a 2027-07-31 deadline, with adjudicators not yet named.

Not yet established: calibration, baseline outperformance, useful lead time, willingness to pay.

What comes before what

The temporal ladder

New medical evidence climbs a ladder of recognition, from first publication through replication and expert consensus to the moment an institution finally acts. We forecast the institutional end of that ladder — guideline, regulatory, and coverage actions — because those are the rungs where something public and dated happens that anyone can check.

For a complete breakdown of all 12 rungs, including the specific observable signals, roles, and validation criteria of each stage, read the Evidence-to-Action Ladder reference.

Three capability tiers, different evaluation criteria

Available now (rungs 1–4)

Refract reconstructs claim trajectories: emergence, wording evolution, evidence attachment, persistence, propagation. The demonstrated capability.

Judged against: reproducibility from public sources, claim-identity maintenance across revisions.

Research program (rungs 5–7)

Tests whether trajectory features forecast rung 5–7 transitions (guideline, regulatory, coverage actions). The protocol below is what the research program will use — it has not yet produced scored forecasts.

Judged against: Brier score, calibration, lead time, baseline comparison. None established yet.

Forecast anatomy

Three things, kept apart

What the record showed, what we estimated from it, and what the authority actually did are three separate things. Keeping them apart is what makes scoring honest — blur any two and you can no longer tell a good call from a lucky one.

Observed NC-assessed Resolved

Proposition design

What makes a question answerable

A question has to meet eight conditions before we will put a number on it. The system checks them at registration, so a question that cannot be settled never becomes a forecast in the first place.

Temporal integrity

Keeping hindsight out

The easiest way to look good at forecasting is to let information from after the fact leak into the estimate. Every probability is pinned to a dated evidence state to make that impossible, and each registration stands on its own rather than borrowing from the ones before it.

Evidence cutoff

All sources included in the forecast are frozen at the evidence cutoff date. Evidence that arrives after this date is excluded. The cutoff is recorded in the registration and cannot be modified.

Immutable registration

Once registered, the forecast — its probability, baseline, sources, and proposition — is locked. No edits, no backfills, no silent corrections. The registration hash is the record. See how to verify one on the <a href="/verification/">verification page</a>.

Version independence

If a source is updated after registration, the original version is preserved in the forecast record. The forecast scores against the original evidence state, not the current one.

When we forecast the same thing twice

The same question can be forecast more than once, at different cutoffs or against different deadlines. Each of those is its own registration, scored on its own, against the evidence that existed when it was made. A later estimate never rescues an earlier one, and an earlier one never gets credit for what a later one got right.

Authority architecture

Nobody grades their own work

We read the evidence and we make the estimate. We do not decide whether we were right. That last step goes to an independent resolver applying a rule written before the outcome was known, because a forecaster who scores their own forecasts is not producing a track record.

Observed NC-assessed Resolved

Reference forecasts

What we have to beat

A probability that looks impressive on its own may be no better than guessing the historical rate. So every forecast is scored next to four alternatives you could have used instead. If it does not beat them, it has not earned anything, however good the raw number looks.

Base-rate

How often has this type of transition occurred historically for similar authorities and similar evidence states? The base-rate is the simplest reference forecast.

Heuristic

Rule-of-thumb estimates derived from observable features: time since last change, source count, evidence trajectory direction. Transparent, auditable, reproducible.

Expert

Structured elicitation from domain experts, calibrated against the same scoring rules. Expert baselines are scored alongside model baselines — no privileged status.

LLM

Language-model estimates under controlled prompting. Scored against the same proper scoring rule as every other baseline. Not privileged, not excluded.

Outcome adjudication

Who decides what happened

A resolver with no stake in the forecast takes the rule we published before the horizon opened and applies it to the source documents. Their job is narrow on purpose: they do not weigh evidence and they do not touch probabilities. They read what the authority published and say whether it meets the rule.

Clear occurrence

The authority enacted the defined transition within the horizon. The outcome is scored as 1.0. The source document that confirms the transition is cited.

Clear non-occurrence

The horizon closed and the transition did not occur. The outcome is scored as 0.0. The absence is confirmed against the same source set used at registration.

Ambiguous

The transition partially occurred, the authority made a related but not identical move, or the evidence is genuinely unclear. The resolver applies the pre-stated resolution rule. If the rule does not resolve it, the forecast is excluded from scoring rather than forced.

Evaluation

How the scoring works

Each registration is scored against the resolved public outcome using a proper scoring rule — one that gives a forecaster no reason to shade a number toward whatever would look better later. Backtests over history and live forecasts going forward are reported separately, because the two prove very different things.

Proper scoring rule

Brier score and log loss. Proper scoring rules ensure that the forecaster's best strategy is to report their true belief — no incentive to shade or inflate.

Calibration

Calibration curves per cohort, per horizon, and per baseline type. Overconfident forecasts and underconfident forecasts are both failures. The calibration record shows which.

Baseline comparison

Every forecast is scored relative to each baseline. A forecast that does not outperform the base-rate has not demonstrated value regardless of its absolute accuracy.

Lead time over rung

Days between NC crossing the actionable threshold and the target transition (or comparator rung event). Reported with precision, recall, calibration, and false-alert burden at that threshold.

The calibration instrument

Forecasts plot against outcomes by reference class as they resolve. It currently holds none.

The denominator

50 forecast slots, loaded empty. Each fills with a registered forecast and the status it resolves to.

Core metric

Lead time, and why it is never reported alone

Any system can produce enormous lead time by calling everything early. What matters is how much warning you get at a level of precision you can actually act on, counting the false alarms it costs you. So lead time is always published alongside those figures, never on its own.

Lead time formula

Lead time = date of target transition − date the forecast first crossed the frozen actionable threshold.

Always reported with

Precision at threshold, recall, calibration, false-alert burden. Not multiplied — always reported together.

Comparator timestamps

Each comparator requires an explicit observation rule: KOL convergence threshold, RWE publication date, procedural signal date, final action date.

Multiple lead-time measures

Lead over KOL convergence, lead over first qualifying RWE, lead over procedural authority signal, lead over final authority action. Different products, not pooled.

Why we never pool one record into another

A label change is not a guideline revision, and a coverage decision is not a formulary move. Doing well at one says nothing about the others, so each kind of decision keeps its own record. Pooling them is how a weak record hides inside a strong one, and we do not allow it.

Classification

Each proposition is tagged with the authority type (regulatory, guideline, payer), the action type (label change, guideline revision, coverage decision), and the evidence context (trial-based, observational, consensus-based).

Separation

Calibration curves and scoring records are computed within each event class. A forecast's demonstrated accuracy in one class does not transfer to another.

Exclusions

When an event class has too few resolved forecasts for meaningful calibration, the class is flagged as underpopulated. Forecasts in that class are still scored, but the calibration record notes the limited sample.

What is public, and what stays yours

The research is public because that is the only way the method can be checked by anyone who did not build it. The work we do with a customer is not — same rules, same baselines, same scoring, but the questions and the context belong to them. One side shows whether the method works at all; the other shows whether it is worth anything for a particular decision.

Public research
Method validation

Historical transitions, scoring records, calibration curves, baseline comparisons. Published at nextconsensus.com/research/.

Scoring records

Prospective and retrospective forecast scores, Brier scores, log loss, calibration by cohort and horizon. Published at nextconsensus.com/evaluation/.

Evidence surfaces

Historical reconstructions test temporal integrity. The example specification tests proposition structure. Neither is presented as prospective performance.

What this tests

Whether the method can define resolvable propositions, freeze evidence state, issue calibrated probabilities, and score outcomes — using historical data and public sources.

Enterprise application
Forecast program

Propositions tailored to the enterprise's evidence context, scored against the same baselines and scoring rules as public research.

Evidence state

Sources, qualifiers, populations, endpoints, and authority claims the forecast depends on — frozen at the evidence cutoff.

Scoring record

What was forecast, what was resolved, how the scoring compared to baselines, and what the calibration record shows.

What this tests

Whether the method produces useful forecasts for real decision contexts — whether the probability is well-calibrated, whether the baselines are competitive, whether the source trail is verifiable.

When we decline to forecast

A number attached to a record too thin to support it is a guess wearing a decimal point. When the public evidence will not carry an estimate, we say so and issue nothing. These are the conditions under which that happens.

Thin record

When the public source record for a proposition is sparse — few relevant studies, no recent guideline updates, limited citation history — the forecast is abstained rather than forced. The gap is stated explicitly.

Ambiguous edits

When a source was edited but the clinical meaning is unclear — a wording change that might or might not affect the transition — the forecast flags the ambiguity. Verification is needed before a probability can be assigned.

Missing context

When the forecast depends on information not in the public record — a jurisdiction-specific guideline, an internal policy, a private source — the dependency is stated. The forecast is incomplete until the gap is filled.

Competing signals

When the evidence record shows contradictory signals — one source supports the transition, another challenges it — both are presented with their authorities and timestamps. The record does not force a resolution.

Method boundary

When the proposition falls outside the system's defined source universe or resolution capability — a domain the protocol does not monitor, or an action type it is not designed to adjudicate — the boundary is stated rather than claimed.

Safety surface

Adverse-event, REMS, boxed-warning, and recall-safety transitions are outside the source universe by design. NextConsensus observes promotional, efficacy, and coverage positions only; nothing in the system ingests or scores safety signals.

Declining is on the record too

A forecaster who quietly skips the hard questions will look better than one who does not. So every decline is published with its reason and date, next to the scoring record. How often we pass is part of what you are judging us on.

Reason

Why the forecast was abstained: thin record, ambiguous edits, missing context, competing signals, or method boundary.

Condition

The specific evidence gap, ambiguity, or boundary that triggered the abstention. Stated precisely, not gestured at.

Impact

What the exclusion means for the scoring record. Abstained forecasts are not scored. They are recorded as abstained with the reason. The abstention rate is part of the method's public performance record.

How we will be wrong

Every forecast is an estimate from the sources available on a given day, and some of them will miss. These are the ways that happens, named in advance so a failure can be attributed rather than explained away after the fact.

Forecast is wrong

A transition was forecast with high probability and did not occur, or vice versa. The scoring record captures this. The calibration record shows whether it is a pattern or an outlier.

Misses a transition

A transition occurred but was not forecast, or was assigned low probability. The attribution record identifies the gap — was it a source outside the monitored universe, a model error, or a genuinely surprising event?

Assessment is ambiguous

Sometimes the evidence record is genuinely unclear. A guideline updates but the recommendation language stays the same. A trial result is mixed. The resolver applies the pre-stated rule; if it does not resolve, the forecast is excluded.

Sources are incomplete

The system is designed to forecast from the public source set. If relevant evidence exists only in private materials, internal databases, or subscription-only sources, the forecast is limited to what can be inspected and cited.

Both directions cost something. A false alarm burns your team's attention on work that did not need doing; a miss leaves you flat-footed on the day it publishes. Any program should measure both within its event class. The method makes that tradeoff visible — it does not make it go away, and anyone claiming otherwise is selling something.

When it does not happen, where did it stop?

Recording a non-event as a plain "no" throws away almost everything useful about it. A guideline that stalled because the evidence never replicated failed for a completely different reason than one the committee never got to. So we record which:

Evidentiary failure (rungs 1–2)

The signal did not replicate or weakened materially.

Recognition failure (rung 3)

Evidence persisted, but expert interpretation did not converge.

Translation failure (rung 4→5)

Experts converged, but the relevant authority did not initiate or advance action.

Procedural delay (rung 5–7)

The transition appeared directionally supported but missed the forecast horizon because of process timing.

Institutional resistance (rung 5–7)

The authority declined to move despite strong upstream evidence or recognition.

Competing transition

A different action occurred than the one forecast — e.g., narrowing instead of expansion.

Ambiguous resolution

The authority changed language, but not enough to meet the pre-registered threshold.

Late occurrence

The event occurred, but after the specified deadline. Still a miss for the original proposition.

Every record keeps the target outcome, how far the transition actually got, and what happened afterward. That is worth far more than a column of yes and no — both for the customer asking why something stalled and for the method learning to read the next one.

Estimating one step at a time

Rather than guessing at the final outcome in one leap, the model estimates the chance of each step happening — and how long a change has already been sitting where it is. Chaining those steps gives the forecast, and it also shows where a transition is most likely to get stuck, which is usually the more useful answer.

Transition Forecast

Probability of reaching target rung by deadline, composed from the chain of rung-to-rung hazards.

Bottleneck Diagnosis

Which rung-to-rung transition is currently limiting downstream movement (e.g., insufficient RWE confirmation, low expert convergence, procedural not-ready).

Interpretability

The chain reveals where a transition is likely to stall. More useful than a single opaque probability.

Conditioning variables

Evidence state, dwell time per rung, authority behavior, event class, disease area, precedent, procedural status.

Method basis

Why the method focuses on date, source, and limits.

These works frame the forecasting burden: evidence records age, guidelines change, and transparent scoring matters more than unsupported certainty.