Voidly's classifier v3.3 works well on countries with many labels (Iran, Russia, China — all at F1 ≈ 1.0). But a quiet failure hides in the long tail: OM, UZ, TN, LY, YE, JO, MA and ten other MENA / former-Soviet states regress 5–29pp F1 compared to v3.1. We tried a regime-cluster fine-tune (v3.4) and it didn't help — documented in that finding.
The root cause: with fewer than 5 confirmed censorship incidents, there is no per-country signal for ANY parametric model to learn from. v3.3 effectively falls back to the global prior — a flat ~16.7% block rate — for every tail country. That is the wrong answer for Qatar (87%), Kuwait (76%), and Bahrain (68%) and the right answer for almost no one.
Instead of asking "what does this country's history say about itself?", we ask "which other countries does this one look like, and what do they say?" That's zero-shot cross-country transfer — fit a regressor from per-country meta-features → empirical block rate on the countries that DO have signal, then project every country (including the 16 tail ones) into that same feature space.
Meta-features we use (39 total after one-hot encoding):
The label is the empirical block_rate over the past 365 days. Historical features sit in a window 365–730 days ago — disjoint from the label window — so the model can't trivially memorise.
We bake off Ridge (alpha 1.0), Ridge-strong (alpha 10), GradientBoosting (max_depth 2 / 3, n_estimators 100 / 200) under LOOCV. gbm_shallow (max_depth=2, n_estimators=100, lr=0.05) wins:
| candidate | LOOCV R² |
|---|---|
| ridge (α=1) | 0.245 |
| ridge_strong (α=10) | 0.313 |
| gbm_shallow (depth 2) | 0.487 |
| gbm (depth 3) | 0.468 |
| gbm_wide (depth 3, 200 trees) | 0.453 |
v3.3 effectively predicts the global prior (~16.7%) for every tail country. Zero-shot uses meta-features. Direct comparison:
| country | empirical | zero-shot pred | v3.3 prior | zero-shot abs err | v3.3 abs err | improvement |
|---|---|---|---|---|---|---|
| TN | 0.66 | 0.49 | 0.17 | 0.17 | 0.49 | +32pp |
| JO | 0.50 | 0.56 | 0.17 | 0.06 | 0.33 | +27pp |
| QA | 0.88 | 0.44 | 0.17 | 0.44 | 0.71 | +27pp |
| UZ | 0.44 | 0.50 | 0.17 | 0.05 | 0.28 | +22pp |
| SY | 0.46 | 0.55 | 0.17 | 0.09 | 0.30 | +20pp |
| LY | 0.42 | 0.48 | 0.17 | 0.06 | 0.25 | +19pp |
| OM | 0.37 | 0.40 | 0.17 | 0.03 | 0.21 | +18pp |
| MA | 0.53 | 0.33 | 0.17 | 0.19 | 0.36 | +16pp |
| YE | 0.55 | 0.33 | 0.17 | 0.22 | 0.38 | +16pp |
| BH | 0.68 | 0.17 | 0.17 | 0.51 | 0.51 | +0.2pp |
| TJ | 0.04 | 0.21 | 0.17 | 0.17 | 0.13 | −4pp |
| TM | 0.30 | 0.12 | 0.17 | 0.18 | 0.13 | −4pp |
| KW | 0.76 | 0.09 | 0.17 | 0.68 | 0.60 | −8pp |
| KG | 0.34 | 0.06 | 0.17 | 0.28 | 0.18 | −10pp |
| DZ | 0.26 | 0.46 | 0.17 | 0.20 | 0.09 | −11pp |
| IQ | 0.20 | 0.51 | 0.17 | 0.31 | 0.04 | −27pp |
10 of 16 tail countries (62.5%) beat v3.3's flat prior in absolute error. Median tail improvement: +16pp.
mean_block_rate_365d (the feature window overlapped the label window). LOOCV R² was 0.9997 — clearly suspicious. Fixed by moving features to a 365–730-day-ago window before promote.
Promote floor was set at: LOOCV R² ≥ 0.30 AND median tail improvement
vs v3.3 ≥ 0pp AND ≥60% of tail beats v3.3 prior in absolute terms.
All three are met (R² 0.487, median improvement +16.2pp, 10/16 = 62.5% beat).
Live at GET /v1/forecast/zero-shot/<cc>.
GET /v1/forecast/zero-shot/<cc> — block-rate point estimate + 90% CI for any country.GET /v1/forecast/zero-shot/info — full sidecar (candidates, LOOCV metrics, tail evaluation, honest caveats).GET /v1/forecast/zero-shot/coverage — all 148 country predictions in one call.