Track record

Every prediction, graded.

When a public run makes a claim about the future, it goes on this page. When reality arrives, the claim is graded — hit or miss — and the grade never comes down. Misses get a post-mortem. This page is the only credential we ask you to check.

Public predictions

Counting begins at launch. The number will be real, even when it's small.

Calibration score

Published after the first ten gradings. Methodology below.

Misses published

All of them

Every miss gets a post-mortem. None are removed.

The ledger Illustrative sample rows — live ledger begins at launch

Run What the engine said What happened Grade
#0007 Port dwell time at the affected terminal exceeds 6 days within 3 weeks — 71% of runs Dwell peaked at 7.2 days in week 2 HIT
#0006 Grid load-shed events stay under 3 for the month — 84% of runs Five load-shed events occurred MISS post-mortem
#0005 ICU diversion risk crosses 20% if admissions rise 18% Grading window open until season end OPEN

When we miss

Post-mortems, in public

A miss is information. Every graded miss gets a short public post-mortem answering three questions: which assumption was wrong, whether the miss was inside or outside the stated confidence interval, and what changed in the model as a result.

A well-calibrated instrument is supposed to miss sometimes — a claim made at 80% confidence should fail roughly one time in five. What the post-mortems watch for is the difference between honest tail events and systematic bias. Both are reported here, because you cannot judge the difference unless we show you all of it.

Post-mortem · run #0006 · sample

Grid load-shed miss: the demand assumption was stale

The run assumed regional demand growth of 2.1% based on the prior year. Actual demand ran 5.8% higher after two data-centre commissionings the schema's demand distribution had no signal for. The miss fell outside the stated 84% interval. Change made: the energy-grid schema now takes a declared large-load pipeline as an input, and the demand prior was re-fitted. The original run remains replayable, unedited.

How grading works

The rules of the ledger

A prediction is eligible for the ledger when it is public, has a stated probability or interval, a defined grading window, and a resolution source named in advance. Grading is binary against the stated claim — no partial credit, no re-scoring, no quiet edits. If a claim turns out to be ungradeable, it is marked void with the reason, and voids are counted and shown.

The calibration score compares stated confidence with observed frequency across all graded claims: when the engine says 80%, does the thing happen about 80% of the time? A perfect score is not the goal — an honest one is.

Full methodology →