Voidly's production 7-day shutdown forecast (forecast
v1, XGBoost + isotonic) reports an ROC AUC of 0.954.
A weakness audit found that number is not real. It comes from a
shuffled train_test_split in
scripts/train-forecast.py — the shuffle scatters
rows of the same country across the train and test folds, so the model is
graded on dates that sit days apart from rows it trained on. The target is
time-autocorrelated, so that leakage hands the model an almost-free score.
This finding is the honest re-evaluation, plus an attempt to fix the model
with genuinely forward-looking features (forecast v2
momentum). The headline up front: under an honest
forward-temporal split, neither v1 nor v2 has predictive signal beyond a
trivial persistence baseline. v2 is not promoted. v1 stays in
production unchanged — its value is calibration and explanation, not lift.
Everything below uses a forward-temporal split only: train
on all country-days up to a cutoff, test on the strictly-future 60-day
window (2026-03-24 to 2026-05-22, 1,260 rows, 198 positive). No shuffling.
This mirrors how the live forecast actually runs — predict tomorrow from
history alone. The comparison baseline is persistence
("predict-yesterday"): forecast[country, t] =
label[country, t-1].
| Model / split | AUC | F1 | Verdict |
|---|---|---|---|
| v1, shuffled stratified split (the reported number) | 0.954 | 0.641 | leaky — not real |
| v2, shuffled stratified split | 0.989 | 0.851 | leaky — not real |
| v1, forward-temporal split (honest) | 0.589 | 0.132 | near chance |
| v2, forward-temporal split (honest) | 0.685 | 0.404 | still loses to persistence |
| persistence baseline (predict-yesterday) | 0.957 | 0.922 | see below — autocorrelation artifact |
v2 beats v1 on the honest split (+9.6pp AUC, +27pp F1) — the momentum and event features do help relative to v1. But v2 still loses to persistence by −27.2pp AUC and −51.9pp F1. The promote gate required v2 to beat persistence by ≥ 8pp F1. It missed by ~60pp. v2 is not promoted.
A persistence baseline scoring 0.92 F1 looks like 7-day forecasting is a
solved problem. It is not. The target target_7day
is a sliding 7-day window: "was there a censorship incident
in the next 7 days". Two adjacent days share 6 of their 7 lookahead days, so
the label is 98.9% autocorrelated day-to-day — across 15,330
country-day rows there are only 172 transitions (a row whose
label differs from yesterday's). Predict-yesterday wins because the windows
overlap by construction, not because it forecasts anything.
This is the exact same autocorrelation that the shuffled split leaks into v1. The shuffle puts day t in train and day t+1 in test; their labels are nearly identical; the model gets graded as if it predicted the future when it merely memorized a neighbor. Strip the leakage and the real forward AUC of v1 collapses from 0.954 to ~0.59.
The only rows where a forecast is non-trivial are transition rows — where the label actually moves. We restricted the holdout to those 31 rows and scored both model and baseline there:
This is the honest core of the finding. On the days that matter — the onset of a shutdown, the lift of a block — the v2 forecast has no skill, and arguably negative skill. Persistence cannot be beaten not because it is strong, but because the signal the model would need to beat it is not in the data.
v2 keeps all 39 v1 features and adds 30 genuinely forward/change-oriented ones across six families, specifically chosen because v1's features are all level features (lags, rolling means) which are persistence:
Feature-gain attribution confirms the model tried to use the new signal — event-anticipation is the #2 family by gain (15%, behind only the v1 base at 27%), and contagion chain is #3 (6%). The features are not dead weight; they carry the lift that moved v2 from AUC 0.59 to 0.69. But 0.69 is still well under the 0.96 persistence wall. The honest conclusion: with the data Voidly currently has, the 7-day censorship target is persistence-dominated — what is blocked stays blocked, and the rare transitions are not anticipated by momentum, calendar proximity, or neighbor-country contagion.
The production forecast is not deleted, and the /v1/forecast/{cc}/7day
endpoint is unchanged. But the honest framing of what it delivers changes:
the 7-day forecast is best understood as a calibrated persistence
signal with explanation — it tells you which countries are in a
sustained blocking regime and surfaces SHAP drivers + conformal intervals.
It does not reliably call the onset of a new shutdown 7 days out.
The headline AUC on the public model-info endpoints is the leaky stratified
number; readers should treat the honest forward-temporal AUC (~0.59 for v1,
~0.69 for v2) as the real predictive-skill figure. The Sentinel alert
lead-time retrospective (79% false-alarm rate, separately published) is the
same story measured from the alerting side.
Feature builder: scripts/build-forecast-v2-features.py.
Trainer + honest evaluation: scripts/train-forecast-v2-momentum.py
(forward-temporal split is the headline; the transition-only block is the
deepest honest cut). Run on Vultr. Artifacts (all unpromoted):
ml-deploy/censorship_forecast_v2_momentum.pkl,
ml-deploy/forecast_v2_momentum_baseline_gate.json
(the headline gate),
ml-deploy/forecast_v2_momentum_summary.json.
Live transparency at /v1/forecast/v2-momentum/info.
Production /v1/forecast/model/info still reports
v1.
Negative results count. The leaky 0.954 was a real bug; replacing it with an honest 0.59 — and publishing the no-promote — is the fix. A forecast you can trust to be honestly weak is more useful than one that is dishonestly strong.
After this finding shipped, a follow-up audit spot-checked the other forecast variants in the Atlas ML stack — multi-horizon, hourly, and per-region — to see whether they inherit the same autocorrelated-label inflation. They do. The honest scope of v1 applies to all of them, and the public surfaces for each now carry the same current-regime-not-onset disclosure.
| Variant | Label | Evaluation split | Verdict |
|---|---|---|---|
Multi-horizon (1d/7d/30d, /v1/forecast/{cc}/multi-horizon) |
target_Nday — the same sliding-window construction; build-forecast-multi-horizon-labels.py even recomputes target_7day to match the legacy autocorrelated label exactly |
shuffled train_test_split + LOCO — no forward-temporal split, no persistence baseline |
Same inflation as v1. The published LOCO AUC 0.91 / 0.88 / 0.84 inherits the leak. The 1d horizon is the least autocorrelated and 30d the most; none are honest onset predictors. |
Per-region (/v1/forecast/region/{slug}) |
None of its own — it is an evidence-volume-weighted mean of the per-country v1 predict_risk() outputs |
inherits v1's evaluation entirely | Inherits v1's exact onset problem (transition-row AUC ~0.33). A regional aggregate of a current-regime signal is still a current-regime signal. |
Hourly (K=6/12/24h, /v1/forecast/{cc}/hourly) |
future-incident-start in the next K hours — a sliding K-hour window, so still autocorrelated hour-to-hour | proper 30-day temporal holdout (train-hourly-forecast.py splits on a date cutoff — no shuffle leak) |
Better off than the others: the headline AUC is not inflated by the shuffle leak. But the sliding K-hour label is still autocorrelated and there is no transition-only / persistence check, so the hourly AUC still over-credits persistence and should not be read as onset skill. |
The honest takeaway is unchanged and now applies stack-wide: every Voidly
forecast endpoint is a current-regime risk signal. The
machine-readable honest metrics live at /v1/forecast/onset-skill
(and are injected into /v1/forecast/v2-momentum/info
as an onset_skill block). The disclosure now
appears on every public forecast surface — /atlas/forecast,
/atlas/forecast/{cc}, the per-country
RiskForecastChart and
MultiHorizonStrip components, the
HighRiskForecast widget, /methodology,
and /atlas/changelog — via a single shared
ForecastHonestyNote component, so no reader
sees a forecast number without the scope caveat next to it.