The scoreboard

Calibration

Every forecast publishes an 80% interval before resolution and is graded against the official first print when it lands. This page is computed from the same records that back log.json and reward.json; the daily pre-registration chain lives in the public records repository.

Scoring methodology v5 (2026-07-10): headline numbers count only witness-verified scores. A run enters the headline when its sealed custody root was externally witnessed — an RFC 3161 timestamp in the witnessed chronology extracted from the public record chain — before the observation, its custody inventory is complete and headline-eligible, and its recorded run time precedes the observation (sub-day ordering trusted only for explicit UTC-offset timestamps). A claimed timestamp alone never enters the headline: scores whose chronology rests on claimed times stay published in log.json, flagged claimed-time-verified, outside the official numbers, alongside unverified and violated legacy runs. CRPS is normalized only by same-series ledger dispersion frozen at target registration; scores without three pre-cutoff ledger observations publish raw CRPS and stay out of normalized means and rewards. The agent-versus-persistence headline is a per-target RAW CRPS ratio against the paired ledger baseline, which needs no scale at all — nothing a forecast authors can move its denominator.

Scored forecasts
8
witness-verified
witness-verified; 201 claimed-time-only and 62 unverified or violated excluded, 915 awaiting resolution
80% interval coverage
38%
witness-verified
3 of 8 witness-verified observed inside the stated interval
Unpaired mean normalized CRPS
2.350
witness-verified
Lower is better; 2 of 8 witness-verified scores have a ledger scale. Mean sharpness 3.47× target scale.
CRPS ratio vs persistence
1.18
witness-verified
Geometric mean of per-target raw CRPS ratios (agent / paired persistence baseline) on 2 matched targets; below 1 beats persistence. 50% agent win rate.
The claimed-time tier
209scores verified by recorded timestamps
75%interval coverage (157 of 209 inside)
0chronology violations

Scores whose runs predate the witnessed custody chain: their recorded timestamps order them before their outcomes, but no independent timestamp proves it, so they sit one tier below the headline — published and flagged in log.json, never deleted. The headline starts at zero by construction and fills as custody-v2 runs resolve against official prints.

The calibration curve

Each score records the forecast distribution evaluated at the official print (its probability integral transform), so coverage is checkable at every stated interval, not just the elicited 80%. A calibrated forecaster tracks the diagonal: above it means intervals are too wide, below it too narrow. The PIT histograms show the same thing distributionally — a calibrated forecaster fills each bin equally; a U shape is overconfidence, a central hump underconfidence. The curve reads coverage off each forecast's materialized distribution, while the headline 80% stat counts stated interval endpoints directly, so the two can differ by a few scores.

0%0%20%20%40%40%60%60%80%80%100%100%Perfect calibration: observed coverage equals stated coverageelicited 80%Claimed-time tier: 28 of 201 outcomes (14%) inside the stated 10% central intervalClaimed-time tier: 57 of 201 outcomes (28%) inside the stated 20% central intervalClaimed-time tier: 79 of 201 outcomes (39%) inside the stated 30% central intervalClaimed-time tier: 96 of 201 outcomes (48%) inside the stated 40% central intervalClaimed-time tier: 112 of 201 outcomes (56%) inside the stated 50% central intervalClaimed-time tier: 128 of 201 outcomes (64%) inside the stated 60% central intervalClaimed-time tier: 137 of 201 outcomes (68%) inside the stated 70% central intervalClaimed-time tier: 149 of 201 outcomes (74%) inside the stated 80% central intervalClaimed-time tier: 175 of 201 outcomes (87%) inside the stated 90% central intervalClaimed-time tier: 180 of 201 outcomes (90%) inside the stated 95% central intervalWitness-verified: 1 of 8 outcomes (13%) inside the stated 10% central intervalWitness-verified: 2 of 8 outcomes (25%) inside the stated 20% central intervalWitness-verified: 2 of 8 outcomes (25%) inside the stated 30% central intervalWitness-verified: 2 of 8 outcomes (25%) inside the stated 40% central intervalWitness-verified: 2 of 8 outcomes (25%) inside the stated 50% central intervalWitness-verified: 3 of 8 outcomes (38%) inside the stated 60% central intervalWitness-verified: 3 of 8 outcomes (38%) inside the stated 70% central intervalWitness-verified: 3 of 8 outcomes (38%) inside the stated 80% central intervalWitness-verified: 6 of 8 outcomes (75%) inside the stated 90% central intervalWitness-verified: 6 of 8 outcomes (75%) inside the stated 95% central intervalstated central intervaloutcomes inside
Witness-verified (n=8)Claimed-time tier (n=201)
PIT — witness-verified
PIT 0.0–0.1: 3 of 8 outcomesPIT 0.1–0.2: 0 of 8 outcomesPIT 0.2–0.3: 1 of 8 outcomesPIT 0.3–0.4: 0 of 8 outcomesPIT 0.4–0.5: 1 of 8 outcomesPIT 0.5–0.6: 1 of 8 outcomesPIT 0.6–0.7: 0 of 8 outcomesPIT 0.7–0.8: 0 of 8 outcomesPIT 0.8–0.9: 0 of 8 outcomesPIT 0.9–1.0: 2 of 8 outcomesUniform reference: a calibrated forecaster fills each bin equally00.51
PIT — claimed-time tier
PIT 0.0–0.1: 34 of 201 outcomesPIT 0.1–0.2: 12 of 201 outcomesPIT 0.2–0.3: 15 of 201 outcomesPIT 0.3–0.4: 21 of 201 outcomesPIT 0.4–0.5: 27 of 201 outcomesPIT 0.5–0.6: 29 of 201 outcomesPIT 0.6–0.7: 18 of 201 outcomesPIT 0.7–0.8: 17 of 201 outcomesPIT 0.8–0.9: 9 of 201 outcomesPIT 0.9–1.0: 19 of 201 outcomesUniform reference: a calibrated forecaster fills each bin equally00.51

Forecasters against the baseline

Per-target raw CRPS ratio against the paired ledger persistence baseline (geometric mean; below 1 beats persistence), lowest first. Unpaired means remain visible for context. The persistence baseline forecasts every target as its last official print with a realized-volatility interval — an agent earns its place by beating it. Forecaster rows score here only when witness-verified; the deterministic baseline is a replayable function of pre-cutoff ledger data and needs no witness of its own. Rows with few scored runs are noisy; read them accordingly.

ForecasterScored / total runsClaimed-time scoredCRPS ratio vs persistencePaired win rateUnpaired mean nCRPS80% coverage
thesis.analystgpt-5.6-sol4 / 4601.1850%2.3500%
brier.time_series_priorpersistence.last_printbaseline3 / 301.60133%
thesis.analystgpt-5.54 / 2406775%
prototype seed0 / 3130
brier-1.controlgpt-5.40 / 30
brier-1.packedgpt-5.40 / 40
Three-agent CPI ensembleCodex recorded agent ensemble0 / 22
scout-2.controlgpt-5-mini0 / 99
brier-1.packedgpt-50 / 99
brier-1.shadowgpt-50 / 3030
thesis.analyst.laddergpt-5.60 / 10
thesis.analystgpt-5.6-luna0 / 50
thesis.analystgpt-5.6-terra0 / 180
thesis.analyst.median3gpt-5.6-terra0 / 60
thesis.analyst.laddergpt-5.50 / 135
thesis.analyst.median3gpt-5.50 / 125
thesis.analyst.laddergpt-5.6-sol0 / 50
thesis.analyst.median3gpt-5.6-sol0 / 60
thesis.analyst.ladder_v2gpt-5.50 / 60
thesis.analyst.ladder_v2gpt-5.6-sol0 / 60
thesis.analyst.ladder_v2gpt-5.6-terra0 / 60
thesis.analyst.median3gpt-5.6-luna0 / 10
thesis.analyst.ladder_v2gpt-5.6-luna0 / 50
UK indicator agent ensembleCodex recorded agent runs0 / 107
Canada/Australia indicator agent ensembleCodex recorded agent runs0 / 99
Euro area/Japan indicator agent ensembleCodex recorded agent runs0 / 88
US near-term public outcomes agentCodex recorded agent run0 / 1616
brier-defense-public-dataCodex recorded source-context synthesis0 / 50
Occupation automation exposure source synthesisCodex recorded source-context synthesis0 / 60
brier-occupation-projectionCodex recorded source-context synthesis0 / 60
BLS Employment ProjectionsBLS 2024-2034 projection release, OEWS-compatible interpolation0 / 60
brier-cps-occupation-fast-proxyCodex recorded source-context synthesis0 / 66
brier-occupation-automation-scenariosCodex recorded source-context synthesis0 / 120
BLS Employment ProjectionsBLS 2024-2034 projection release0 / 60
brier-occupation-wage-pressureCodex recorded source-context synthesis0 / 1320
BLS OEWS current tableMay 2025 annual 10th percentile wage carry-forward baseline0 / 220
BLS OEWS current tableMay 2025 annual 25th percentile wage carry-forward baseline0 / 220
BLS OEWS current tableMay 2025 annual median wage carry-forward baseline0 / 220
BLS OEWS current tableMay 2025 annual mean wage carry-forward baseline0 / 220
BLS OEWS current tableMay 2025 annual 75th percentile wage carry-forward baseline0 / 220
BLS OEWS current tableMay 2025 annual 90th percentile wage carry-forward baseline0 / 220
Global near-term indicator source synthesisCodex recorded source-context synthesis0 / 88
thesis.analystclaude-fable-50 / 2117
thesis.analystgpt-5-codex0 / 540
thesis.analystdamped_log_trend_v1 + Brier component check0 / 20

Evaluation splits

Rows are split by resolutionDate, not run order. Training code may use only rows whose official resolution was known before the evaluation cutoff.

train
0 / 181
Resolved before 2026-07-01.
validation
5 / 92
Resolved from 2026-07-01 through 2026-12-31.
test
0 / 0
Resolved on or after 2027-01-01.
unresolved
0 / 915
Not eligible for reward until the first-print resolver posts a fact.

Latest resolutions

Both published chronology tiers appear here. Only rows marked witnessed — custody root externally witnessed before the observation — count toward the headline numbers above; claimed rows rest on recorded timestamps alone.

ForecastPredictedObservedIn intervalElicitationChronologyNormalized CRPS
Euro area flash HICP annual inflation, July 20262.6%2.9%yesinterval seededclaimed
Canada GDP by industry, May 2026+0.1%+0.3%yesinterval seededclaimed
US initial claims, week ending July 25212k197knointerval seededwitnessed1.856
US initial claims, week ending July 25208k197knointerval seededclaimed1.274
US Continued Claims, July 18 Week1.8M1.8Mnointerval seededwitnessed
US core PCE MoM June 2026+0.3%+0.1%nointerval seededclaimed
US nursing-home total nurse staffing HPRD, July 2026 first print3.923.86yesinterval seededwitnessed
US nursing home occupancy, July 202679.82%80.45%nointerval seededwitnessed
Australia CPI inflation, Jun 20264.4%3.8%yesinterval seededclaimed
US durable goods orders MoM, June 2026+1.8%+0.3%yesinterval seededwitnessed
US Durable Goods Shipments MoM, Jun 2026+0.6%+0.7%yesinterval seededwitnessed
Australia employment change, June 202618k76knointerval seededwitnessed

Full history: the Thesis Log · method: why forecasting is the harness