The full narrative is unavailable.
The original evidence text could not be loaded. Use the source record below; availability of linked files is checked separately.
v3.3 uses a single 0.5 decision threshold across 131 countries. Computing per-country F1-optimal thresholds via precision-recall sweep lifts median F1 +4pp (mean +5.4pp) for 73 countries with sufficient labels. 41 countries improve by ≥3pp. CG (+35.7pp), OM (+31.7pp), ZW (+20.1pp) are the biggest…
v3.3 uses a single 0.5 decision threshold across 131 countries. Computing per-country F1-optimal thresholds via precision-recall sweep lifts median F1 +4pp (mean +5.4pp) for 73 countries with sufficient labels. 41 countries improve by ≥3pp. CG (+35.7pp), OM (+31.7pp), ZW (+20.1pp) are the biggest movers. Live in /v1/classifier/score/{cc} as label_per_country + threshold_used + threshold_source fields.
The original evidence text could not be loaded. Use the source record below; availability of linked files is checked separately.
Original source links and labels. Linked responses may have changed since publication.