Kubernetes Beginner About 20 min +160 XP

Order Scoring Pods Keep Restarting

Both order-scoring worker pods die within seconds of every start, and the orders.score backlog keeps climbing.

Stage 3 of 5 in Kubernetes Night Shift.

Briefing

The order-scoring workers were redeployed at 09:23 as part of a manifest cleanup.

Since then both pods have restarted six times, each run lasting only a few seconds, and the orders.score backlog has passed a thousand messages.

The worker image is the same 14 January build that has run all week.

Dev points out that the data team retrained the scoring models on Monday and suspects the bigger models are the problem, so he would like to pin the old model files.

The fix is yours to find: hints, the recovery checks and the debrief unlock inside the incident.

How it plays

  1. 01

    Get paged

    The alert fires and the clock starts. Read the page and the briefing.

  2. 02

    Investigate

    Work in a simulated shell with realistic output: logs, configs, services.

  3. 03

    Fix it

    Change the system the way you would in production. Hints are there if you get stuck.

  4. 04

    Prove it

    Automated checks verify the recovery, then the debrief explains what happened.

Your pager is ready.

Free, instant, and it works on your phone. No signup: start as a guest and save your progress later.