We trained three LightGBM quantile regressors (alpha=0.05, 0.50, 0.95) on the same 15,351-row country-day forecast dataset that powers /v1/forecast/{cc}/7day. The goal was to ship a journalist-grade p5..p95 band alongside the existing point estimate so reporters citing Voidly could quote a confidence interval rather than a single number.

We are not promoting this model to a live endpoint. The reason is instructive and worth documenting publicly: the model is correct, the metric is correct, the gate is correct, and the conclusion is "the data does not admit a well-calibrated low-quantile estimate."

What we built

LOCO coverage (the honest number)

QuantileNominalEmpiricalErrorWithin ±5pp?
p55.0%81.3%+76.3ppNO
p5050.0%90.5%+40.5ppNO
p9595.0%98.0%+3.0ppYES

Only the upper bound (p95) is well-calibrated. The lower and median quantiles massively over-cover.

Why this happens

The target target_sum_7day is zero-inflated: 80% of country-day windows have exactly zero shutdown-positive days. The positive rate of the binary target_7day is 5.2%; the mean of the continuous version is 0.23 (windows that do contain shutdowns rarely cover the entire week).

Quantile regression on this distribution converges to:

With q05 = 0 nearly always, P(y ≤ q05) = P(y = 0) ≈ 0.80, so the empirical coverage of the p5 prediction is 80%, not 5%. This is not a model bug — a calibrated p5 on this distribution would require negative quantiles, which is nonsensical for "share of days censored".

CQR (the standard post-hoc fix) shifts the band by an additive constant. We measured that constant on a held-out calibration fold per LOCO iteration; the learned shift was 0.002 — essentially zero, because the model is already at the edge of the achievable distribution.

What we'd need to ship this

Three viable paths forward:

  1. Switch to a Beta or Tweedie likelihood that natively handles zero-inflation. LightGBM does not support these out of the box; we'd need an external GLM or a custom objective.
  2. Two-stage hurdle model: stage 1 predicts P(any shutdown), stage 2 predicts the conditional intensity given a shutdown occurred. This is a separate, more involved build.
  3. Frame the question differently: instead of "share of days censored", forecast "expected shutdown-days conditional on at least one shutdown" — a censored regression target on a much smaller sample (n ≈ 800).

None of those is a quick add-on, so we are publishing the negative result rather than papering over it. The existing point estimate at /v1/forecast/{cc}/7day remains the authoritative shutdown forecast; its 90% conformal interval (delivered by ACI on every response) already gives a band — derived from residual conformal inference, not quantile regression, and properly calibrated to nominal coverage.

Top features (incidental finding)

Even though the calibration story is negative, the feature importance from the median model is informative — it tells us what the model thinks drives shutdown intensity, not just occurrence:

  1. block_rate_roll14_mean (65.9% gain) — the trailing 14-day average block rate dominates.
  2. block_rate_roll7_std (11.8%) — volatility in recent block rates.
  3. block_rate_roll7_mean (9.3%) — 7-day trailing average.
  4. block_rate_lag7 (4.2%) — 7-day-lagged value.
  5. week_of_year (1.7%) — modest seasonality.

This is consistent with the v3.3 classifier: rolling block-rate features explain the vast majority of gain. Event features (elections, protests, GDELT unrest) contribute <1% of total gain — the model is essentially a smoother over recent block rates with a small seasonal correction.

Artefacts

The model file is preserved so future work (hurdle models, Beta likelihoods) has a baseline to beat.

Methodology citations