A Temporal Fusion Transformer (TFT, Lim et al. 2021,
arXiv:1912.09363)
trained on the same 38-feature panel that powers our XGBoost
per-horizon forecasts. The TFT predicts a 30-day p10/p50/p90
trajectory in a single attention-based forward pass, replacing
what would otherwise be three independent gradient-boosted
models and a Gaussian conformal wrapper. Live at
GET /v1/forecast/{cc}/tft;
metadata at
GET /v1/forecast/tft/info.
21 watched countries × 731 days = 15,351 country-day rows. Each sample uses a 30-day encoder window over 33 observed reals (lag and rolling block-rate stats, OONI anomaly counts, IODA alerts, probe consensus flags, GDELT protest/conflict signals) plus 5 known calendar covariates (day of week, week of year, month, weekend flag, Friday flag). Static categoricals are the country ID and its tier-1-through-5 risk-tier label.
We use the canonical TFT building blocks — variable selection networks, gated residual blocks, an LSTM encoder, an interpretable multi-head attention decoder — with hidden size 8 and a single attention head (small enough to converge on CPU in the 60-minute budget the directive required). The loss is a quantile loss over (0.1, 0.5, 0.9), which is the textbook way to get a calibrated forecast interval out of a probabilistic regressor without a separate conformal step.
We benchmark against the existing multi-horizon XGBoost stack (1-day / 7-day / 30-day separate models, each with its own isotonic calibrator). The bar to promote was set deliberately high: LOCO median AUC must beat the 7-day baseline (0.9236) by at least 0.03, with p10/p90 empirical coverage inside ±10pp of the 80% nominal band.
Numbers will be filled in once the training run completes;
we publish the LOCO table inline in
/v1/forecast/tft/info regardless of
the promotion outcome.
honest_caveats array so callers
don't treat the interval as well-calibrated.
Even when the TFT does not strictly beat the boosted stack on
median LOCO AUC, the per-horizon attention curve and the
unified probabilistic output make it a useful second opinion
when the XGBoost models disagree. The endpoint is shipped as a
research artifact, not a replacement — the production
forecast still flows through
/v1/forecast/{cc}/7day.