Voidly's production 7-day shutdown forecast (forecast v1, XGBoost + isotonic) reports an ROC AUC of 0.954. A weakness audit found that number is not real. It comes from a shuffled train_test_split in scripts/train-forecast.py — the shuffle scatters rows of the same country across the train and test folds, so the model is graded on dates that sit days apart from rows it trained on. The target is time-autocorrelated, so that leakage hands the model an almost-free score.

This finding is the honest re-evaluation, plus an attempt to fix the model with genuinely forward-looking features (forecast v2 momentum). The headline up front: under an honest forward-temporal split, neither v1 nor v2 has predictive signal beyond a trivial persistence baseline. v2 is not promoted. v1 stays in production unchanged — its value is calibration and explanation, not lift.

The honest evaluation rule

Everything below uses a forward-temporal split only: train on all country-days up to a cutoff, test on the strictly-future 60-day window (2026-03-24 to 2026-05-22, 1,260 rows, 198 positive). No shuffling. This mirrors how the live forecast actually runs — predict tomorrow from history alone. The comparison baseline is persistence ("predict-yesterday"): forecast[country, t] = label[country, t-1].

The numbers

Model / splitAUCF1Verdict
v1, shuffled stratified split (the reported number)0.9540.641leaky — not real
v2, shuffled stratified split0.9890.851leaky — not real
v1, forward-temporal split (honest)0.5890.132near chance
v2, forward-temporal split (honest)0.6850.404still loses to persistence
persistence baseline (predict-yesterday)0.9570.922see below — autocorrelation artifact

v2 beats v1 on the honest split (+9.6pp AUC, +27pp F1) — the momentum and event features do help relative to v1. But v2 still loses to persistence by −27.2pp AUC and −51.9pp F1. The promote gate required v2 to beat persistence by ≥ 8pp F1. It missed by ~60pp. v2 is not promoted.

Why persistence scores 0.92 — and why that is NOT good news

A persistence baseline scoring 0.92 F1 looks like 7-day forecasting is a solved problem. It is not. The target target_7day is a sliding 7-day window: "was there a censorship incident in the next 7 days". Two adjacent days share 6 of their 7 lookahead days, so the label is 98.9% autocorrelated day-to-day — across 15,330 country-day rows there are only 172 transitions (a row whose label differs from yesterday's). Predict-yesterday wins because the windows overlap by construction, not because it forecasts anything.

This is the exact same autocorrelation that the shuffled split leaks into v1. The shuffle puts day t in train and day t+1 in test; their labels are nearly identical; the model gets graded as if it predicted the future when it merely memorized a neighbor. Strip the leakage and the real forward AUC of v1 collapses from 0.954 to ~0.59.

The transition-only test — the deepest honest cut

The only rows where a forecast is non-trivial are transition rows — where the label actually moves. We restricted the holdout to those 31 rows and scored both model and baseline there:

This is the honest core of the finding. On the days that matter — the onset of a shutdown, the lift of a block — the v2 forecast has no skill, and arguably negative skill. Persistence cannot be beaten not because it is strong, but because the signal the model would need to beat it is not in the data.

What v2 added (and why it still was not enough)

v2 keeps all 39 v1 features and adds 30 genuinely forward/change-oriented ones across six families, specifically chosen because v1's features are all level features (lags, rolling means) which are persistence:

Feature-gain attribution confirms the model tried to use the new signal — event-anticipation is the #2 family by gain (15%, behind only the v1 base at 27%), and contagion chain is #3 (6%). The features are not dead weight; they carry the lift that moved v2 from AUC 0.59 to 0.69. But 0.69 is still well under the 0.96 persistence wall. The honest conclusion: with the data Voidly currently has, the 7-day censorship target is persistence-dominated — what is blocked stays blocked, and the rare transitions are not anticipated by momentum, calendar proximity, or neighbor-country contagion.

Promote gate — fails, as it should

What this means for the product

The production forecast is not deleted, and the /v1/forecast/{cc}/7day endpoint is unchanged. But the honest framing of what it delivers changes: the 7-day forecast is best understood as a calibrated persistence signal with explanation — it tells you which countries are in a sustained blocking regime and surfaces SHAP drivers + conformal intervals. It does not reliably call the onset of a new shutdown 7 days out. The headline AUC on the public model-info endpoints is the leaky stratified number; readers should treat the honest forward-temporal AUC (~0.59 for v1, ~0.69 for v2) as the real predictive-skill figure. The Sentinel alert lead-time retrospective (79% false-alarm rate, separately published) is the same story measured from the alerting side.

What could actually move this

  1. Change the target. Predict incident onset (a 0→1 transition) rather than "incident anywhere in a sliding 7-day window". Onset is not autocorrelated; persistence cannot game it; the model would be measured on the thing that matters. The cost is a far rarer positive class (~172 events) — this needs careful class handling.
  2. Higher-frequency leading signals. Momentum on daily-aggregated OONI is too coarse. Hourly probe data, BGP churn, and IODA sub-daily connectivity may carry onset signal that day-level block-rate deltas wash out.
  3. More labeled transitions. 172 transitions across 21 countries and 2 years is a tiny learning signal. Historical backfill of confirmed shutdown onset/recovery dates would help more than any model architecture change.

Reproducibility

Feature builder: scripts/build-forecast-v2-features.py. Trainer + honest evaluation: scripts/train-forecast-v2-momentum.py (forward-temporal split is the headline; the transition-only block is the deepest honest cut). Run on Vultr. Artifacts (all unpromoted): ml-deploy/censorship_forecast_v2_momentum.pkl, ml-deploy/forecast_v2_momentum_baseline_gate.json (the headline gate), ml-deploy/forecast_v2_momentum_summary.json. Live transparency at /v1/forecast/v2-momentum/info. Production /v1/forecast/model/info still reports v1.

Negative results count. The leaky 0.954 was a real bug; replacing it with an honest 0.59 — and publishing the no-promote — is the fix. A forecast you can trust to be honestly weak is more useful than one that is dishonestly strong.

Addendum (2026-05-22) — the other forecast variants share the same inflation

After this finding shipped, a follow-up audit spot-checked the other forecast variants in the Atlas ML stack — multi-horizon, hourly, and per-region — to see whether they inherit the same autocorrelated-label inflation. They do. The honest scope of v1 applies to all of them, and the public surfaces for each now carry the same current-regime-not-onset disclosure.

VariantLabelEvaluation splitVerdict
Multi-horizon (1d/7d/30d, /v1/forecast/{cc}/multi-horizon) target_Nday — the same sliding-window construction; build-forecast-multi-horizon-labels.py even recomputes target_7day to match the legacy autocorrelated label exactly shuffled train_test_split + LOCO — no forward-temporal split, no persistence baseline Same inflation as v1. The published LOCO AUC 0.91 / 0.88 / 0.84 inherits the leak. The 1d horizon is the least autocorrelated and 30d the most; none are honest onset predictors.
Per-region (/v1/forecast/region/{slug}) None of its own — it is an evidence-volume-weighted mean of the per-country v1 predict_risk() outputs inherits v1's evaluation entirely Inherits v1's exact onset problem (transition-row AUC ~0.33). A regional aggregate of a current-regime signal is still a current-regime signal.
Hourly (K=6/12/24h, /v1/forecast/{cc}/hourly) future-incident-start in the next K hours — a sliding K-hour window, so still autocorrelated hour-to-hour proper 30-day temporal holdout (train-hourly-forecast.py splits on a date cutoff — no shuffle leak) Better off than the others: the headline AUC is not inflated by the shuffle leak. But the sliding K-hour label is still autocorrelated and there is no transition-only / persistence check, so the hourly AUC still over-credits persistence and should not be read as onset skill.

The honest takeaway is unchanged and now applies stack-wide: every Voidly forecast endpoint is a current-regime risk signal. The machine-readable honest metrics live at /v1/forecast/onset-skill (and are injected into /v1/forecast/v2-momentum/info as an onset_skill block). The disclosure now appears on every public forecast surface — /atlas/forecast, /atlas/forecast/{cc}, the per-country RiskForecastChart and MultiHorizonStrip components, the HighRiskForecast widget, /methodology, and /atlas/changelog — via a single shared ForecastHonestyNote component, so no reader sees a forecast number without the scope caveat next to it.