On 2026-05-21 we shipped an isotonic refit that collapsed the
forecast's headline calibration drift from +56.45pp
to ~0pp. That looks great on the global track record at
/v1/atlas/prediction-track-record — but
“global ~0pp” is fully consistent with individual countries still
drifting in opposite directions and silently cancelling each other out
in the aggregate.
Iran can be well-calibrated while Egypt's slice is +20pp under-predicting and Venezuela's is −15pp over-predicting. Aggregate to a single number and the picture looks healthy. Slice it per country and we see the truth.
Every day at 05:00 UTC,
scripts/build-per-country-calibration-drift.py
walks the top-50 most-active forecast countries and computes:
Any country with drift exceeding ±15pp (and at least 15 predictions
in the window) gets flagged, the sidecar at
/opt/voidly-ai/ml-deploy/calibration_drift_by_country.json
is updated, and a calibration_drift event
fires into the existing CenAlerts pipeline with 24h dedup.
GET /v1/sentinel/calibration-drift —
full table, sorted by absolute drift desc. Supports
?status=drifting|ok|insufficient_data
and ?min_n=N filters.GET /v1/sentinel/calibration-drift/{cc}
— single-country slice.GET /v1/sentinel/calibration-drift/info
— methodology + honest caveats.
Each per-country row carries the rebuilt observed labels (so we don't
ding the model for IODA disruption rows it wasn't trained to predict),
the median operational threshold actually emitted in the window, and a
flag that toggles to drift_alert: true
when the ±15pp gate is crossed. Operators can subscribe to the
calibration_drift event in CenAlerts and
get pinged the moment a country starts diverging.
status: insufficient_data
and don't fire alerts.abs_drift_pp
tends to overstate true drift when the country is bouncing between
blocked and unblocked.A forecast model that is honest globally but silently broken on Egypt is worse than a model that is openly broken everywhere — the former gets cited as authoritative, while the actual users in Egypt see wildly wrong numbers. Daily per-country auditing makes that failure mode visible and actionable: when a country crosses ±15pp, the alert fires, the sidecar updates, and the endpoint can be cited as proof that we're watching.