Briefing
The API dashboards went flat twenty minutes ago, and the alert rules built on them went quiet at the same moment, which is its own kind of outage.
The API service is healthy, and its metrics exporter answers when you ask it directly from the host.
The monitoring configuration was edited this afternoon as part of a port clean-up on this host.
Dev thinks the metrics library upgrade in the last API release changed the exposition format and Prometheus is rejecting it.
The fix is yours to find: hints, the recovery checks and the debrief unlock inside the incident.