Voidly
Sentinel · 30-day backtest

Forecasts meet outcomes.

Compare forecasts with outcomes. Calibration measures probability accuracy; it does not establish shutdown-onset skill.

Updated every 30 min · last refresh Oct 3, 2026 · CC BY 4.0 · Binned JSON · Raw outcomes

Reliability diagram

0.000.000.250.250.500.500.750.751.001.00perfectn=448n=26n=48n=48n=42n=56n=75n=101n=52n=4Predicted probability (bin mean)Observed positive rate

Each point is one prediction bin. X axis is the mean predicted probability inside the bin; Y axis is the fraction of those forecasts where the real outcome actually happened. Perfect calibration is the diagonal line — points above the line mean the model UNDER-estimates risk; points below mean it OVER-estimates.

Bubble area scales with bin count · Red = model under-estimated · Blue = model over-estimated

Brier score
0.310
lower is better
Calibration MAE
0.369
0 = perfect
Accuracy
53.9%
900 evaluated
F1 (binary 0.5)
0.09
P=0.07 R=0.15

Per-bin breakdown

BinPredicted meanObserved rateΔn
[0.0, 0.1)0.0320.261+0.229448
[0.1, 0.2)0.1370.000-0.13726
[0.2, 0.3)0.2720.125-0.14748
[0.3, 0.4)0.3420.104-0.23848
[0.4, 0.5)0.4550.071-0.38342
[0.5, 0.6)0.5510.071-0.48056
[0.6, 0.7)0.6490.040-0.60975
[0.7, 0.8)0.7560.050-0.707101
[0.8, 0.9)0.8370.000-0.83752
[0.9, 1.0)0.9410.000-0.9414

Δ = observed − predicted. The 0.1 bin holds 448 of the 900 forecasts — this is where most action happens, and where the May 20 isotonic recalibration was aimed. See /sentinel/calibration for the time-series view of how this gap evolves day over day.

Per-country backtest (worst Brier first, n ≥ 5)

CountryBrierAccuracyPRnPos rate
BangladeshBD——1.000.0330—
BrazilBR——0.201.0030—
BelarusBY——0.00—30—
ChinaCN——0.00—30—
CubaCU——0.00—30—
EgyptEG——0.041.0030—
ERER——0.00—30—
EthiopiaET——0.00—30—
IndonesiaID——0.330.0530—
IndiaIN———0.0030—
IranIR——0.000.0030—
North KoreaKP——0.00—30—
KazakhstanKZ——0.201.0030—
LebanonLB————30—
MyanmarMM——0.00—30—
MalaysiaMY——1.000.0530—
NigeriaNG——0.00—30—
NicaraguaNI————30—
PhilippinesPH——0.00—30—
PakistanPK——0.00—30—

Countries where the forecast is currently performing worst — useful for targeting feature engineering or seeking expert review.

How to read these numbers

  • Brier score — mean squared error between predicted probability and actual 0/1 outcome. Lower is better. Less than 0.10 is excellent; 0.10-0.30 is OK; above 0.30 is concerning.
  • Calibration MAE — average gap between predicted-mean and observed-rate across bins. 0.00 means the model's probabilities are exactly right on average.
  • Reliability diagram — the visual version of calibration MAE. Bubble size = bin sample count.
  • F1 (P + R) — binary classification metrics at the 0.5 threshold. Useful when downstream decisions are binary (alert / no-alert).
  • The May 20, 2026 isotonic recalibration targeted the 0.1 bin specifically — see the recalibration finding.
About this view

When the Sentinel model says “5% risk,” does the real outcome actually happen ~5% of the time? Below is the answer: 900 live (predicted, observed) pairs from the last 30 days, binned into a standard reliability diagram.

Calibration and onset skill