Kubernetes Advanced About 40 min +320 XP

Sensor Agent Crash-Looping on Every Node

The endpoint sensor on all five fleet-east nodes has been crash-looping.

Briefing

Quillon Security runs its Sentry endpoint sensor on every node of the fleet-east cluster as the sentry-agent DaemonSet, and at 09:04 all five sensor pods started crash-looping inside the same minute.

Agent 7.5.0 went out on Tuesday with new named-pipe hooks, so the obvious suspect is already being named in the channel.

One node, fleet-east-4, has shown SchedulingDisabled.

Inspired by a real outage: External Technical Root Cause Analysis — Channel File 291 (CrowdStrike, 2024). Names and details are fictionalized.

The fix is yours to find: hints, the recovery checks and the debrief unlock inside the incident.

How it plays

  1. 01

    Get paged

    The alert fires and the clock starts. Read the page and the briefing.

  2. 02

    Investigate

    Work in a simulated shell with realistic output: logs, configs, services.

  3. 03

    Fix it

    Change the system the way you would in production. Hints are there if you get stuck.

  4. 04

    Prove it

    Automated checks verify the recovery, then the debrief explains what happened.

Your pager is ready.

Free, instant, and it works on your phone. No signup: start as a guest and save your progress later.