What the global track record was hiding

On 2026-05-21 we shipped an isotonic refit that collapsed the forecast's headline calibration drift from +56.45pp to ~0pp. That looks great on the global track record at /v1/atlas/prediction-track-record — but “global ~0pp” is fully consistent with individual countries still drifting in opposite directions and silently cancelling each other out in the aggregate.

Iran can be well-calibrated while Egypt's slice is +20pp under-predicting and Venezuela's is −15pp over-predicting. Aggregate to a single number and the picture looks healthy. Slice it per country and we see the truth.

What this monitor does

Every day at 05:00 UTC, scripts/build-per-country-calibration-drift.py walks the top-50 most-active forecast countries and computes:

Any country with drift exceeding ±15pp (and at least 15 predictions in the window) gets flagged, the sidecar at /opt/voidly-ai/ml-deploy/calibration_drift_by_country.json is updated, and a calibration_drift event fires into the existing CenAlerts pipeline with 24h dedup.

Endpoints

How it reads

Each per-country row carries the rebuilt observed labels (so we don't ding the model for IODA disruption rows it wasn't trained to predict), the median operational threshold actually emitted in the window, and a flag that toggles to drift_alert: true when the ±15pp gate is crossed. Operators can subscribe to the calibration_drift event in CenAlerts and get pinged the moment a country starts diverging.

Honest caveats

Why this matters

A forecast model that is honest globally but silently broken on Egypt is worse than a model that is openly broken everywhere — the former gets cited as authoritative, while the actual users in Egypt see wildly wrong numbers. Daily per-country auditing makes that failure mode visible and actionable: when a country crosses ±15pp, the alert fires, the sidecar updates, and the endpoint can be cited as proof that we're watching.