Does the uncertainty hold?
Check observed coverage against the model’s uncertainty. Calibration does not establish shutdown-onset skill.
Empirical coverage — 90-day rolling
The blue line is the actual fraction of forecasts where the real outcome landed inside the 90% conformal interval. The green dashed line is the nominal target (0.90). If the blue stays close to the green, the model is well calibrated.
Blue: empirical coverage · Dashed green: nominal 0.90 target · Orange circles: drift alerts
Live forecast accuracy (prod_rolling, 30-day window)
Brier < 0.10 is good, > 0.30 is concerning. Calibration MAE < 0.05 means predicted-probabilities track observed-rates closely. See /sentinel/backtest for the actual reliability diagram (predicted-mean vs observed-rate scatter) and /methodology#validation for the full evaluation methodology + 3-split honest baselines.
Last 14 days
| Date | Coverage | q90 | n holdout | Drift? |
|---|---|---|---|---|
| Oct 3 | 90.3% | 0.244 | 2,180 | — |
| Oct 2 | 90.3% | 0.244 | 2,180 | — |
| Oct 1 | 90.3% | 0.244 | 2,180 | — |
| Sep 30 | 90.3% | 0.244 | 2,180 | — |
| Sep 29 | 90.3% | 0.244 | 2,180 | — |
| Sep 28 | 90.3% | 0.244 | 2,180 | — |
| Sep 27 | 90.3% | 0.244 | 2,180 | — |
| Sep 26 | 90.0% | 0.252 | 2,158 | — |
| Sep 25 | 90.0% | 0.252 | 2,158 | — |
| Sep 24 | 90.0% | 0.252 | 2,158 | — |
| Sep 23 | 90.0% | 0.252 | 2,158 | ⚠️ |
| Sep 22 | 90.0% | 0.252 | 2,158 | ⚠️ |
| Sep 21 | 90.0% | 0.252 | 2,158 | ⚠️ |
| Sep 20 | 90.0% | 0.252 | 2,158 | ⚠️ |
What features the model actually uses
Sklearn feature_importances_ on the underlying XGBoost. 39 features total. Top-3 sum: 0.615 · Top-5: 0.658 · Top-10: 0.747. Healthy distribution — no single feature dominates the model.
- 1.recent_shutdown55.6%
- 2.block_rate_roll30_mean3.5%
- 3.critical_incident_7d2.5%
- 4.week_of_year2.4%
- 5.verified_signals_7d1.9%
- 6.month1.9%
- 7.high_importance_event1.8%
- 8.election_in_7days1.8%
- 9.blocked_count_roll7_mean1.8%
- 10.block_rate_roll14_mean1.7%
- 11.gdelt_unrest_30d1.4%
- 12.incident_count_7d1.4%
Interpretation: The forecast model's top feature is gdelt_unrest_30d (0.25) — protest + conflict signals from the GDELT 1.0 global news feed. recent_shutdown, block_rate rolling means, and incident counts follow. risk_tier — the leaky country-level encoding that dominated our older classifier at 85% — contributes only ~2% here. Healthy distribution; no single feature dominates.
Raw JSON: /v1/sentinel/feature-importance
Related
- /methodology#validation — 3-split honest baselines (LOCO 0.91 AUC vs stratified 0.98)
- /v1/sentinel/accuracy — full evaluation JSON, updated nightly
- /atlas/elections — see the forecast in action (90-day upcoming elections)
About this view
Every Sentinel shutdown forecast ships with a 90% conformal interval. This page tracks how often the real outcome lands inside that interval — the closer to 90%, the more honest the model. Data lives at /v1/sentinel/calibration/history and updates every 24h.
Model history & calibration scope
/v1/forecast/{cc}/7day response under aci_alpha + aci.* fields. Full ACI methodology →