Voidly Atlas now exposes 30+ live ML endpoints —
forecast (7d, multi-horizon, TFT, zero-shot, region, platform, domain,
ASN-GNN, duration, trajectory, hourly), classifier (v3.3 +
feature-importance + robustness), anomaly (DBSCAN, IsolationForest,
per-domain HDBSCAN drift), per-measurement classifier, plus the
sentinel transparency family (accuracy, feature-importance, global
heatmap, calibration-drift, prediction-track-record). For a journalist
who just wants to know “is the model live and working right now?”
that's 30+ separate curl commands.
This dashboard collapses them into a single answer.
What it measures
Every day at 04:00 UTC a cron on Vultr runs
scripts/build-serving-reliability.py.
For each endpoint it:
- fires 10 probes over a ~2-minute window directly
at
127.0.0.1:5002 (forecast_api.py
upstream)
- computes HTTP availability (% probes returning 2xx)
- computes p50 + p95 latency over the successful
probes (note: noisy at n=10 — intentional)
- parses the JSON body and runs a top-level-keys superset
check against the expected schema
- computes model staleness from the most recent of
either a declared
trained_at field
in the sidecar or the sidecar file mtime
- records whether the response surfaces an
honest_caveats field
(Voidly's standard for ML transparency)
- for forecast/sentinel groups, pulls the latest global calibration
drift from
forecast_calibration_refit.json
Promote criteria
- ≥ 25 endpoints probed successfully (one of the published
summary numbers)
- The dashboard is honest about endpoints that 500'd or
timed out — failures are not hidden, they're
surfaced in
summary.worst_availability
and per-endpoint errors dict.
Honest caveats
- 10-probe sample window is small. p95 latency at
n=10 is dominated by tail noise. Treat it as a smoke signal, not
a contract.
- Doesn't measure tail-of-distribution failures.
A daily snapshot misses 5-minute blips during peak load.
- Probes hit localhost upstream. CF gateway / DNS
/ TLS issues are not detected here — that's a
separate health check.
- Schema check is a fingerprint, not a full
JSON-schema validation. We confirm the expected top-level keys
are present (case-insensitive), not their nested shape.
- MCP server endpoints are excluded. They run on
a separate Node.js runtime and have their own publish process.
Live at
GET /v1/atlas/serving-reliability — full sidecar JSON
GET /v1/atlas/serving-reliability?group=forecast — filter by group
GET /v1/atlas/serving-reliability/info — metadata + caveats