Research methodology and status
What's available, what's in validation, what's not yet proven
NextConsensus separates demonstrated capability from research hypothesis. Claim-trajectory reconstruction is available now. Claim prioritization is in validation. Authority-transition forecasting is a research program — an earlier proof cycle appeared to show signal, but the result was invalid (truncated negative controls). The first record, FF-001, is registered against a 2027-07-31 deadline, with adjudicators not yet named.
Jump to a section
Where things stand
Capability status
Available now
Claim-trajectory reconstruction. Refract reconstructs how medical claims emerge, change wording, acquire evidence, persist or disappear, and spread across influential sources — from dated public revision history.
Every event is traceable to a public source with a reproducible hash.
In validation
Claim prioritization. Signs of movement — gaining stronger support, becoming harder to dismiss, spreading across influential sources, losing qualifiers, appearing in regulator or guideline language — are being tested for their ability to rank what deserves attention now.
Validated through the 19-drug study and buyer-discovery interviews.
Research program
Authority-transition forecasting. Tests whether trajectory features predict institutional transitions. An earlier proof cycle was invalid (truncated negative controls). The first record, FF-001, is registered against a 2027-07-31 deadline, with adjudicators not yet named.
Not yet established: calibration, baseline outperformance, useful lead time, willingness to pay.
What comes before what
The temporal ladder
New medical evidence climbs a ladder of recognition, from first publication through replication and expert consensus to the moment an institution finally acts. We forecast the institutional end of that ladder — guideline, regulatory, and coverage actions — because those are the rungs where something public and dated happens that anyone can check.
For a complete breakdown of all 12 rungs, including the specific observable signals, roles, and validation criteria of each stage, read the Evidence-to-Action Ladder reference.
Three capability tiers, different evaluation criteria
Available now (rungs 1–4)
Refract reconstructs claim trajectories: emergence, wording evolution, evidence attachment, persistence, propagation. The demonstrated capability.
Judged against: reproducibility from public sources, claim-identity maintenance across revisions.
Research program (rungs 5–7)
Tests whether trajectory features forecast rung 5–7 transitions (guideline, regulatory, coverage actions). The protocol below is what the research program will use — it has not yet produced scored forecasts.
Judged against: Brier score, calibration, lead time, baseline comparison. None established yet.
Forecast anatomy
Three things, kept apart
What the record showed, what we estimated from it, and what the authority actually did are three separate things. Keeping them apart is what makes scoring honest — blur any two and you can no longer tell a good call from a lucky one.
The public record as it existed at the evidence cutoff. Sources, versions, timestamps, and gaps — frozen at the moment the forecast is registered. This is what the forecast is about.
A probability estimate that a defined medical authority will make a specified public transition within a stated horizon. Registered once, immutable after registration, scored against the independently adjudicated outcome.
The adjudicated result: did the transition occur, not occur, or was it judged ambiguous? Resolved by an independent party with no stake in the forecast. Published with the scoring record.
Proposition design
What makes a question answerable
A question has to meet eight conditions before we will put a number on it. The system checks them at registration, so a question that cannot be settled never becomes a forecast in the first place.
“Will GLP-1s get broader coverage?”
↓ the same question, made resolvable ↓
Will CMS issue a national coverage determination extending coverage of Therapy X to the second-line population on or before 2027-06-30?
The sentence carries four conditions. Four more are record fields.
Temporal integrity
Keeping hindsight out
The easiest way to look good at forecasting is to let information from after the fact leak into the estimate. Every probability is pinned to a dated evidence state to make that impossible, and each registration stands on its own rather than borrowing from the ones before it.
All sources included in the forecast are frozen at the evidence cutoff date. Evidence that arrives after this date is excluded. The cutoff is recorded in the registration and cannot be modified.
Once registered, the forecast — its probability, baseline, sources, and proposition — is locked. No edits, no backfills, no silent corrections. The registration hash is the record. See how to verify one on the <a href="/verification/">verification page</a>.
If a source is updated after registration, the original version is preserved in the forecast record. The forecast scores against the original evidence state, not the current one.
When we forecast the same thing twice
The same question can be forecast more than once, at different cutoffs or against different deadlines. Each of those is its own registration, scored on its own, against the evidence that existed when it was made. A later estimate never rescues an earlier one, and an earlier one never gets credit for what a later one got right.
Reference forecasts
What we have to beat
A probability that looks impressive on its own may be no better than guessing the historical rate. So every forecast is scored next to four alternatives you could have used instead. If it does not beat them, it has not earned anything, however good the raw number looks.
How often has this type of transition occurred historically for similar authorities and similar evidence states? The base-rate is the simplest reference forecast.
Rule-of-thumb estimates derived from observable features: time since last change, source count, evidence trajectory direction. Transparent, auditable, reproducible.
Structured elicitation from domain experts, calibrated against the same scoring rules. Expert baselines are scored alongside model baselines — no privileged status.
Language-model estimates under controlled prompting. Scored against the same proper scoring rule as every other baseline. Not privileged, not excluded.
Outcome adjudication
Who decides what happened
A resolver with no stake in the forecast takes the rule we published before the horizon opened and applies it to the source documents. Their job is narrow on purpose: they do not weigh evidence and they do not touch probabilities. They read what the authority published and say whether it meets the rule.
The authority enacted the defined transition within the horizon. The outcome is scored as 1.0. The source document that confirms the transition is cited.
The horizon closed and the transition did not occur. The outcome is scored as 0.0. The absence is confirmed against the same source set used at registration.
The transition partially occurred, the authority made a related but not identical move, or the evidence is genuinely unclear. The resolver applies the pre-stated resolution rule. If the rule does not resolve it, the forecast is excluded from scoring rather than forced.
Evaluation
How the scoring works
Each registration is scored against the resolved public outcome using a proper scoring rule — one that gives a forecaster no reason to shade a number toward whatever would look better later. Backtests over history and live forecasts going forward are reported separately, because the two prove very different things.
Brier score and log loss. Proper scoring rules ensure that the forecaster's best strategy is to report their true belief — no incentive to shade or inflate.
Calibration curves per cohort, per horizon, and per baseline type. Overconfident forecasts and underconfident forecasts are both failures. The calibration record shows which.
Every forecast is scored relative to each baseline. A forecast that does not outperform the base-rate has not demonstrated value regardless of its absolute accuracy.
Days between NC crossing the actionable threshold and the target transition (or comparator rung event). Reported with precision, recall, calibration, and false-alert burden at that threshold.
The calibration instrument
Forecasts plot against outcomes by reference class as they resolve. It currently holds none.
The denominator
50 forecast slots, loaded empty. Each fills with a registered forecast and the status it resolves to.
Core metric
Lead time, and why it is never reported alone
Any system can produce enormous lead time by calling everything early. What matters is how much warning you get at a level of precision you can actually act on, counting the false alarms it costs you. So lead time is always published alongside those figures, never on its own.
Lead time = date of target transition − date the forecast first crossed the frozen actionable threshold.
Precision at threshold, recall, calibration, false-alert burden. Not multiplied — always reported together.
Each comparator requires an explicit observation rule: KOL convergence threshold, RWE publication date, procedural signal date, final action date.
Lead over KOL convergence, lead over first qualifying RWE, lead over procedural authority signal, lead over final authority action. Different products, not pooled.
Why we never pool one record into another
A label change is not a guideline revision, and a coverage decision is not a formulary move. Doing well at one says nothing about the others, so each kind of decision keeps its own record. Pooling them is how a weak record hides inside a strong one, and we do not allow it.
Each proposition is tagged with the authority type (regulatory, guideline, payer), the action type (label change, guideline revision, coverage decision), and the evidence context (trial-based, observational, consensus-based).
Calibration curves and scoring records are computed within each event class. A forecast's demonstrated accuracy in one class does not transfer to another.
When an event class has too few resolved forecasts for meaningful calibration, the class is flagged as underpopulated. Forecasts in that class are still scored, but the calibration record notes the limited sample.
What is public, and what stays yours
The research is public because that is the only way the method can be checked by anyone who did not build it. The work we do with a customer is not — same rules, same baselines, same scoring, but the questions and the context belong to them. One side shows whether the method works at all; the other shows whether it is worth anything for a particular decision.
Historical transitions, scoring records, calibration curves, baseline comparisons. Published at nextconsensus.com/research/.
Prospective and retrospective forecast scores, Brier scores, log loss, calibration by cohort and horizon. Published at nextconsensus.com/evaluation/.
Historical reconstructions test temporal integrity. The example specification tests proposition structure. Neither is presented as prospective performance.
Whether the method can define resolvable propositions, freeze evidence state, issue calibrated probabilities, and score outcomes — using historical data and public sources.
Propositions tailored to the enterprise's evidence context, scored against the same baselines and scoring rules as public research.
Sources, qualifiers, populations, endpoints, and authority claims the forecast depends on — frozen at the evidence cutoff.
What was forecast, what was resolved, how the scoring compared to baselines, and what the calibration record shows.
Whether the method produces useful forecasts for real decision contexts — whether the probability is well-calibrated, whether the baselines are competitive, whether the source trail is verifiable.
When we decline to forecast
A number attached to a record too thin to support it is a guess wearing a decimal point. When the public evidence will not carry an estimate, we say so and issue nothing. These are the conditions under which that happens.
When the public source record for a proposition is sparse — few relevant studies, no recent guideline updates, limited citation history — the forecast is abstained rather than forced. The gap is stated explicitly.
When a source was edited but the clinical meaning is unclear — a wording change that might or might not affect the transition — the forecast flags the ambiguity. Verification is needed before a probability can be assigned.
When the forecast depends on information not in the public record — a jurisdiction-specific guideline, an internal policy, a private source — the dependency is stated. The forecast is incomplete until the gap is filled.
When the evidence record shows contradictory signals — one source supports the transition, another challenges it — both are presented with their authorities and timestamps. The record does not force a resolution.
When the proposition falls outside the system's defined source universe or resolution capability — a domain the protocol does not monitor, or an action type it is not designed to adjudicate — the boundary is stated rather than claimed.
Adverse-event, REMS, boxed-warning, and recall-safety transitions are outside the source universe by design. NextConsensus observes promotional, efficacy, and coverage positions only; nothing in the system ingests or scores safety signals.
Declining is on the record too
A forecaster who quietly skips the hard questions will look better than one who does not. So every decline is published with its reason and date, next to the scoring record. How often we pass is part of what you are judging us on.
Why the forecast was abstained: thin record, ambiguous edits, missing context, competing signals, or method boundary.
The specific evidence gap, ambiguity, or boundary that triggered the abstention. Stated precisely, not gestured at.
What the exclusion means for the scoring record. Abstained forecasts are not scored. They are recorded as abstained with the reason. The abstention rate is part of the method's public performance record.
How we will be wrong
Every forecast is an estimate from the sources available on a given day, and some of them will miss. These are the ways that happens, named in advance so a failure can be attributed rather than explained away after the fact.
A transition was forecast with high probability and did not occur, or vice versa. The scoring record captures this. The calibration record shows whether it is a pattern or an outlier.
A transition occurred but was not forecast, or was assigned low probability. The attribution record identifies the gap — was it a source outside the monitored universe, a model error, or a genuinely surprising event?
Sometimes the evidence record is genuinely unclear. A guideline updates but the recommendation language stays the same. A trial result is mixed. The resolver applies the pre-stated rule; if it does not resolve, the forecast is excluded.
The system is designed to forecast from the public source set. If relevant evidence exists only in private materials, internal databases, or subscription-only sources, the forecast is limited to what can be inspected and cited.
Both directions cost something. A false alarm burns your team's attention on work that did not need doing; a miss leaves you flat-footed on the day it publishes. Any program should measure both within its event class. The method makes that tradeoff visible — it does not make it go away, and anyone claiming otherwise is selling something.
When it does not happen, where did it stop?
Recording a non-event as a plain "no" throws away almost everything useful about it. A guideline that stalled because the evidence never replicated failed for a completely different reason than one the committee never got to. So we record which:
The signal did not replicate or weakened materially.
Evidence persisted, but expert interpretation did not converge.
Experts converged, but the relevant authority did not initiate or advance action.
The transition appeared directionally supported but missed the forecast horizon because of process timing.
The authority declined to move despite strong upstream evidence or recognition.
A different action occurred than the one forecast — e.g., narrowing instead of expansion.
The authority changed language, but not enough to meet the pre-registered threshold.
The event occurred, but after the specified deadline. Still a miss for the original proposition.
Every record keeps the target outcome, how far the transition actually got, and what happened afterward. That is worth far more than a column of yes and no — both for the customer asking why something stalled and for the method learning to read the next one.
Estimating one step at a time
Rather than guessing at the final outcome in one leap, the model estimates the chance of each step happening — and how long a change has already been sitting where it is. Chaining those steps gives the forecast, and it also shows where a transition is most likely to get stuck, which is usually the more useful answer.
Probability of reaching target rung by deadline, composed from the chain of rung-to-rung hazards.
Which rung-to-rung transition is currently limiting downstream movement (e.g., insufficient RWE confirmation, low expert convergence, procedural not-ready).
The chain reveals where a transition is likely to stall. More useful than a single opaque probability.
Evidence state, dwell time per rung, authority behavior, event class, disease area, precedent, procedural status.
Method basis
Why the method focuses on date, source, and limits.
These works frame the forecasting burden: evidence records age, guidelines change, and transparent scoring matters more than unsupported certainty.
- How quickly do systematic reviews go out of date? A survival analysis
Shows that review currency varies by topic and can decay before normal review schedules catch it.
- Validity of the Agency for Healthcare Research and Quality clinical practice guidelines: how quickly do guidelines become outdated?
Shows that guideline validity changes over time and should be reassessed against new evidence.
- GRADE: an emerging consensus on rating quality of evidence and strength of recommendations
Separates evidence quality from the strength of recommendations and makes uncertainty explicit.
- The PRISMA 2020 statement: an updated guideline for reporting systematic reviews
Supports transparent reporting of search, selection, appraisal, synthesis, and update methods.