Kubernetes Intermediate About 35 min +220 XP

API Pods Restart Across the Rollout

Every api replica is in CrashLoopBackOff and the Deployment has shown 0 of 3 available for ten minutes.

Stage 4 of 4 in Platform Recovery Ladder.

Briefing

The api Deployment moved to the new cluster ten minutes ago with a manifest the platform team tidied up on the way.

Since then all three pods have been restarting every minute or two, and the Deployment has not shown a single available replica.

The image, myapp:v1, is the same build that ran for weeks on the old hosts, and its logs look like a normal boot every time.

Dev is convinced the new nodes are short on memory.

The fix is yours to find: hints, the recovery checks and the debrief unlock inside the incident.

How it plays

  1. 01

    Get paged

    The alert fires and the clock starts. Read the page and the briefing.

  2. 02

    Investigate

    Work in a simulated shell with realistic output: logs, configs, services.

  3. 03

    Fix it

    Change the system the way you would in production. Hints are there if you get stuck.

  4. 04

    Prove it

    Automated checks verify the recovery, then the debrief explains what happened.

Your pager is ready.

Free, instant, and it works on your phone. No signup: start as a guest and save your progress later.