Observability Intermediate About 30 min +230 XP

API Metrics Disappear from Prometheus

API dashboards have gone flat: Prometheus reports its api scrape target down while the exporter answers from the host.

Stage 2 of 3 in Observability Blackout.

Briefing

The API dashboards went flat twenty minutes ago, and the alert rules built on them went quiet at the same moment, which is its own kind of outage.

The API service is healthy, and its metrics exporter answers when you ask it directly from the host.

The monitoring configuration was edited this afternoon as part of a port clean-up on this host.

Dev thinks the metrics library upgrade in the last API release changed the exposition format and Prometheus is rejecting it.

The fix is yours to find: hints, the recovery checks and the debrief unlock inside the incident.

How it plays

  1. 01

    Get paged

    The alert fires and the clock starts. Read the page and the briefing.

  2. 02

    Investigate

    Work in a simulated shell with realistic output: logs, configs, services.

  3. 03

    Fix it

    Change the system the way you would in production. Hints are there if you get stuck.

  4. 04

    Prove it

    Automated checks verify the recovery, then the debrief explains what happened.

Your pager is ready.

Free, instant, and it works on your phone. No signup: start as a guest and save your progress later.