Per-category censorship classifiers (NEWS / ANON / GRP / PORN / COMT / MMED / SRCH)

v3.3 answers is this country-day censoring something? The per-method classifier (also shipped 2026-05-21) answers HOW are they blocking it — DNS, TCP, HTTP, or TLS? This finding answers the journalist-gold question that sits between them: WHAT are they targeting?

The distinction matters editorially. A regime that blocks just ANON (Tor / VPNs / Lantern / Psiphon) is suppressing circumvention — that's a tell of internal political pressure, not free-press repression. A regime that blocks just NEWS (BBC / NYT / Guardian / Substack) is muzzling journalism. A regime that blocks SRCH (Google / DuckDuckGo) is taking a much heavier hand. The 7 promoted per-category classifiers expose those signals at the API.

Approach

Each classifier uses the same 16-feature v3.3 input (anomaly_rate, measurement_count, spike_magnitude, weekday/month/weekend, probe-block-rate, neighbor-contagion features, etc.) trained with class-balanced XGBoost. The category label is positive iff the country-day had ≥1 evidence row whose blocked domain matches that Citizen Lab category. Only ~18 distinct domains in the evidence table have domain_category populated, so we back-fill via a 37-domain hand-curated map covering NEWS (bbc, nytimes, guardian, ...), ANON (tor, protonvpn, psiphon, lantern, ...), COMT (signal, telegram, whatsapp, ...), and the rest. Negatives are category-agnostic (same negative row used for every category) so we don't leak "wrong-category" signal into the negative pool.

Results: 7 of 7 categories with data promoted

CategoryPos labelsStrat AUCLOCO median AUCPromote path
NEWS1760.9420.897alt-AUC
ANON1940.9210.866alt-AUC
GRP2870.9130.858alt-AUC
PORN850.9770.957alt-AUC
COMT2320.9330.856alt-AUC
MMED2300.9350.849alt-AUC (bonus)
SRCH1200.9710.924alt-AUC (bonus)
POLR0——SKIP
RELR0——SKIP

5 of 7 primary categories cleared (needed 4) — family promote passes. POLR (political opposition) and RELR (religion) are honestly skipped because zero rows in our evidence table currently match those Citizen Lab categories. The endpoint returns available: false with a transparent reason rather than a hallucinated probability.

Top countries per category (training-time)

Promote gate

Floors are deliberately softer than v3.3's gate because per-category labels are ~5–7% positive rate (vs v3.3's 26%), so F1 at threshold 0.5 punishes any high-precision-low-recall model. We use either:

All 7 promoted categories cleared the alt-AUC path with margin (lowest LOCO AUC was COMT at 0.856).

Honest caveats

Endpoints