Voidly

Atlas model registry

Know what
the model means.

The target. The version. The evidence.
A useful score starts with all three.

Inspect the models
Different models. Different questions.
Observed evidenceClassifier

Incident classification
and disruption signals

Past OONI signalsEvent forecast

A country-relative event
within the next seven days

This is a target map, not a comparison of accuracy.

Source request

Loaded metadata describes an inspected model bundle. It does not by itself prove which model every prediction route serves.

01 / Evidence classificationIntrospection loaded

Censorship classifier.

v3.3

LOCO mean F10.711
Mean F1 · well-sampled countries0.630

Cross-country evaluation. These values are not per-prediction probabilities or future-time accuracy.

794 of 1,116 positive training labels (71.1%) are country-days whose only incident is an IODA `disruption` row — network outages, not confirmed censorship. The same class was excluded from forecast labels in 2026-05 as ~94% noise. So this model is substantially trained to detect DISRUPTION, and its F1 should be read as such. Corpus is frozen at 2026-05-21; relabelling is a pending decision, not an oversight.

Algorithm
GradientBoostingClassifier
Model trained
2026-05-21T03:01:46.793987+00:00
Record generated
Not supplied
LOCO countries
127
Well-sampled countries
61
Evaluation, training set and all source caveats
Source headline
LOCO mean F1 across all countries
LOCO median F1 · distribution caveat
0.870
Countries with perfect F1
46
Stratified F1 · sanity check
0.729
Stratified AUC · sanity check
0.899
Training samples
4,237
Positive samples
1,116
Training countries
131

The headline LOCO MEDIAN F1 (0.870) is dominated by many small-sample countries scoring a perfect 1.0 on a handful of days; the MEAN F1 (0.711) is the honest single number. Censorship-heavy, high-volume countries score materially lower (CN ~0.29 on n=95, BY ~0.21 on n=84, AZ ~0.11 on n=65). Cite the mean — or the specific per-country number — for hard countries, not the median. The model is CLEAN (no label leakage); this is a distribution caveat, not an accuracy retraction.

HONESTY: forward-temporal holdout (train past, test future) gives AUC 0.669 / F1 0.474 vs the random-split AUC 0.895 / F1 0.725. v3.3 generalizes across COUNTRIES (LOCO F1 0.87) but DEGRADES across TIME (delta AUC -0.226) — which is why it is retrained weekly. Do not read the random-split number as forward-deployment accuracy. See /atlas/findings/classifier-v3.3-temporal-generalization-2026-05.

confidence is

stratified cross-validation F1, NOT a calibrated per-prediction probability

correct use

Cross-country censorship-risk classification. Per-country accuracy varies widely — 16 MENA / former-Soviet countries (OM, UZ, TN, LY, YE, JO, MA, …) regress 5-29pp on sparse neighbor-pair overlap. See evaluation.honest_evaluation + evaluation.per_country for the full distribution.

headline metric

LOCO median F1 0.87 (cross-country); honest MEAN F1 0.71; well-sampled (n>=30 countries) mean ~0.63

model status

CLEAN — no label leakage (ML_LEAKAGE_AUDIT.md); this caveat is about the score distribution, not leakage.

training label composition

794 of 1,116 positive training labels (71.1%) are country-days whose only incident is an IODA `disruption` row — network outages, not confirmed censorship. The same class was excluded from forecast labels in 2026-05 as ~94% noise. So this model is substantially trained to detect DISRUPTION, and its F1 should be read as such. Corpus is frozen at 2026-05-21; relabelling is a pending decision, not an oversight.

Source feature names (16)
  • anomaly_rate
  • measurement_count
  • spike_magnitude
  • day_of_week
  • month
  • is_weekend
  • rate_count_interaction
  • probe_block_rate
  • probe_node_count
  • probe_avg_confidence
  • probe_agreement
  • rate_spike_interaction
  • high_evidence
  • neighbor_block_rate_7d
  • neighbor_incident_count_7d
  • neighbor_max_anomaly_7d
02 / OONI-event forecastIntrospection loaded

Seven-day event signal.

honest_forecast_v1

Rolling-origin PR-AUC0.297
Rolling-origin ROC AUC0.817

Paired precision-recall and ranking measures from temporal validation. They cannot be compared directly with classifier F1.

A country-relative OONI measurement event is not necessarily a national internet shutdown. This record does not establish an exact onset date, a legacy SHAP explanation or a calibrated uncertainty interval.

Forecast horizon
7 days
Source generated
2026-10-03T06:18:58.380418+00:00
Source countries
60
Recommended source threshold
0.650
Target, validation and all source caveats

independent OONI aggregation-API ground truth, country-relative z-score event label, strict temporal rolling-origin CV — see scripts/honest-forecast-backtest.py

The target is any event day in T+1 through T+7. An event requires an anomaly rate above the country’s strictly past baseline by the stated z threshold, at or above the absolute rate floor, and enough measurements. The standard-deviation floor limits unstable z scores. Stable blocking can produce a low change estimate; a low event estimate is not proof of accessibility.

Validation method
strict temporal rolling-origin CV, 9 folds, on independent OONI ground truth — see honest-forecast-backtest.py
Country baseline window
28 days
Event z threshold
2.00
Event-rate floor
0.080
Standard-deviation floor
0.010
Minimum event-day volume
50
Minimum input volume
30
Mean validation precision
0.309
Mean validation recall
0.426
Mean validation Brier score
0.162
LOCO mean AUC
0.730
Recent-event baseline PR-AUC
0.160
Country climatology baseline PR-AUC
0.215

Validation precision and recall use thresholds selected from each fold’s training data. The recommended source threshold is a separate deployment setting. The recent-event baseline counts event days in the seven-day input window; it is not the legacy “predict yesterday” baseline.

Performance estimate is the rolling-origin backtest (AUC ~0.82, PR-AUC ~0.29) — NOT train-set metrics. The deployed model is retrained on all data, so its train-set scores are meaningless.

Uses only OONI measurement signals (anomaly/confirmed rates, volumes, trajectories). No political-event / election / protest features yet — that is the next obvious lift.

Cross-country transfer (LOCO) is weaker for persistently-censored countries (SA/AE/EG/KZ AUC ~0.6) — the model needs some history of a country to forecast it well.

PR-AUC ~0.29 at a 10% base rate is a useful early-warning signal (precision ~0.31, recall ~0.42 at the F1-optimal threshold) — not a precise oracle.

Another target

Shutdown-risk country ranking is a separate model family. Its country-ranking skill and weak within-country timing are not the event forecast’s validation scores.

Keep the record

History, with its limits.

Read registry history ↗

The entries below preserve the older registry’s model history and links. Dated status claims do not establish today’s serving state.

Classifier v3 → v3.3 · May 2026 history

v3 removed v2’s label-derived country-tier feature. v3.1 expanded the data 13.5× to 4,237 samples, including 1,116 positives across 131 countries. v3.2 tested geographic contagion and was held back. v3.3 used three regime-similarity-weighted cross-country contagion features alongside 13 base features; the archived promotion date is May 21, 2026.

Algorithm
GradientBoostingClassifier
Historical stratified F1
v3 0.46 → v3.1 0.67 → v3.3 0.73
Historical LOCO median F1
v3.1 0.82 → v3.3 0.87; small-sample inflation applies
v3.3 LOCO mean F1
0.711; 0.63 across 61 countries with n≥30
Country distribution
46 small-sample countries at perfect F1; hard high-volume CN 0.29, BY 0.21, AZ 0.11
Egypt recovery
v3.1 F1 0.55 → v3.3 F1 0.73 (+18 percentage points)
Regression caveat
16 MENA / former-Soviet countries regressed versus v3.1 with sparse neighbor-pair overlap
Classifier v2 · retired reference
Version / algorithm
v2 / GradientBoostingClassifier
Training date
2026-02-10
Retirement recorded
2026-05-21; superseded by v3.1
Inflated stratified F1 / ROC AUC
99.8% / 1.000
Feature count
39
Label-derived top feature
country_risk_tier, 85% importance

The original entry contains conflicting retired/serving descriptions. Its historical performance is compromised by label leakage, and this archive does not verify that a prediction endpoint still serves v2.

Legacy Sentinel XGBoost + isotonic · history and health

The older registry describes a seven-day watched-country risk model with SHAP drivers and conformal intervals. It is distinct from the OONI-event record above. Its training date is recorded as April 17, 2026; isotonic recalibration was recorded on May 20, 2026.

Algorithm
XGBoost classifier + sklearn IsotonicRegression
Archived random/held-out AUC / F1
0.980 / 0.795; not forward-deployment accuracy
Archived post-refit Brier
0.223, in-sample
Historical refit changes
Brier 0.59 → 0.22; MAE 0.60 → 0.00, in-sample
Watched-country gate in that account
30 censorship-heavy countries

The sliding-window target is highly autocorrelated. These historical holdout and fit metrics do not establish shutdown onset skill. Read the onset audit.

Current health response · separate source

Health-reported model version
v1@2026-09-27T02:11:03.341768
Model loaded at
2026-10-03T14:39:32.072617+00:00
Reported conformal coverage
90.3%
Reported q90
0.244
Source flags stale inputs
No
Source flags unknown freshness
No
Source evaluated freshness
Yes

Coverage and q90 belong to this health response. They are not uncertainty bounds for honest_forecast_v1. Load time is not a training date or a measurement date. Zone-less source times below are retained literally; no timezone is inferred.

Source-provided freshness flags and dates
InputLatest source recordAge / limit (hours)Freshness
evidence2026-10-03T16:43:19.248044Z0.02 / 12.00Within source threshold
incident_evidence2026-10-03T12:00:09.6117854.74 / 12.00Within source threshold
incidents2026-10-03T12:00:09.6117854.74 / 48.00Within source threshold
probe_metrics2026-10-03T16:44:22.203466+00:000.01 / 2.00Within source threshold

Health feature keys

  • recent_shutdown
  • block_rate_roll30_mean
  • critical_incident_7d
  • week_of_year
  • verified_signals_7d
IsolationForest anomaly detector · retired reference
Version
v1
Algorithm
IsolationForest
Trained
2026-02-10
Deprecated
2026-05-08
Archived state
Frozen; 63 MB pickle

This was the pre-v2 unsupervised baseline. The old registry identifies the archived artifact as anomaly_detector_v1.pkl.deprecated.20260508, kept for historical comparison. Its supersession account cited roughly 2,640 labeled incidents and preferred labeled classification over threshold-based anomaly detection. That is the historical rationale, not a fresh evaluation of all operating points.

Match the test to the claim.

Random split

Sanity check.

Stratified k-fold can leak temporal and country patterns. It does not establish deployment performance.

Country holdout

Geographic transfer.

LOCO trains on the other countries. Show the distribution and well-sampled subset alongside the average.

Future holdout

Temporal transfer.

Train before a cutoff and test after it. A different target or split is a different evaluation.