The problem: 16 countries that no per-country model can learn

Voidly's classifier v3.3 works well on countries with many labels (Iran, Russia, China — all at F1 ≈ 1.0). But a quiet failure hides in the long tail: OM, UZ, TN, LY, YE, JO, MA and ten other MENA / former-Soviet states regress 5–29pp F1 compared to v3.1. We tried a regime-cluster fine-tune (v3.4) and it didn't help — documented in that finding.

The root cause: with fewer than 5 confirmed censorship incidents, there is no per-country signal for ANY parametric model to learn from. v3.3 effectively falls back to the global prior — a flat ~16.7% block rate — for every tail country. That is the wrong answer for Qatar (87%), Kuwait (76%), and Bahrain (68%) and the right answer for almost no one.

The fix: project countries into a shared embedding space

Instead of asking "what does this country's history say about itself?", we ask "which other countries does this one look like, and what do they say?" That's zero-shot cross-country transfer — fit a regressor from per-country meta-features → empirical block rate on the countries that DO have signal, then project every country (including the 16 tail ones) into that same feature space.

Meta-features we use (39 total after one-hot encoding):

The label is the empirical block_rate over the past 365 days. Historical features sit in a window 365–730 days ago — disjoint from the label window — so the model can't trivially memorise.

Validation: leave-one-country-out cross-validation

We bake off Ridge (alpha 1.0), Ridge-strong (alpha 10), GradientBoosting (max_depth 2 / 3, n_estimators 100 / 200) under LOOCV. gbm_shallow (max_depth=2, n_estimators=100, lr=0.05) wins:

candidateLOOCV R²
ridge (α=1)0.245
ridge_strong (α=10)0.313
gbm_shallow (depth 2)0.487
gbm (depth 3)0.468
gbm_wide (depth 3, 200 trees)0.453

Tail country uplift vs v3.3's flat prior

v3.3 effectively predicts the global prior (~16.7%) for every tail country. Zero-shot uses meta-features. Direct comparison:

countryempiricalzero-shot predv3.3 priorzero-shot abs errv3.3 abs errimprovement
TN0.660.490.170.170.49+32pp
JO0.500.560.170.060.33+27pp
QA0.880.440.170.440.71+27pp
UZ0.440.500.170.050.28+22pp
SY0.460.550.170.090.30+20pp
LY0.420.480.170.060.25+19pp
OM0.370.400.170.030.21+18pp
MA0.530.330.170.190.36+16pp
YE0.550.330.170.220.38+16pp
BH0.680.170.170.510.51+0.2pp
TJ0.040.210.170.170.13−4pp
TM0.300.120.170.180.13−4pp
KW0.760.090.170.680.60−8pp
KG0.340.060.170.280.18−10pp
DZ0.260.460.170.200.09−11pp
IQ0.200.510.170.310.04−27pp

10 of 16 tail countries (62.5%) beat v3.3's flat prior in absolute error. Median tail improvement: +16pp.

Honest caveats

Promoted

Promote floor was set at: LOOCV R² ≥ 0.30 AND median tail improvement vs v3.3 ≥ 0pp AND ≥60% of tail beats v3.3 prior in absolute terms. All three are met (R² 0.487, median improvement +16.2pp, 10/16 = 62.5% beat). Live at GET /v1/forecast/zero-shot/<cc>.

Endpoints