Storefront Stuck in Maintenance Mode
Every Northstar Shops storefront page answers 503 with the maintenance banner, and checkout has taken no orders for 20 minutes.
Incident catalog
61 simulated incidents across Linux, Networking, Docker, Kubernetes, Security, Git, CI/CD, Web, Observability and Cloud. Each one is a broken system you investigate in a real-feeling shell, fix, and prove recovered. Nothing to install.
17 incidents
Every Northstar Shops storefront page answers 503 with the maintenance banner, and checkout has taken no orders for 20 minutes.
The nightly backup logs a finished archive every run, yet BackupFailure has paged the night shift three nights running.
The capacity forecast projects the rebuilt cache host's disk full within a week at the cache log's usual growth rate.
The 02:00 content sync never published, and product pages are still showing yesterday's prices.
Three operators report that ll output on the ops host differs from what the handoff runbooks show.
The internal incident notes service is down, leaving responders to coordinate the outage from chat scrollback.
Eight jobs are stuck in the build queue, none has reached the compile step in 40 minutes, and deploys are frozen.
QueueWorkerCritical has paged 42 times in the last hour, although the queue workers report zero lag.
The nightly orders database backup exits successfully but leaves nothing in /backups.
The report API does not answer, and the report dashboard shows errors ahead of Finance's 10:00 run.
Batch exports fail with 'no space left on device', and the batch host's root filesystem is at 100%.
The metrics API will not stay up after its 09:18 restart, and the dashboards that query it show no data.
Host CPU has sat above 95% for ten minutes, and storefront response times are climbing with it.
forgeline.dev answers 500 on the web UI and API, and its readiness probe fails from every region.
The process list shows one order-router process near 98% CPU while the other seven sit between 2% and 5%.
Depot scan uploads are being refused.
The Larder delivery marketplace has accepted no stock feed from Greengage.
5 incidents
Storefront pages load again, but cart, search and checkout calls through /api fail for every customer.
Package installs and API calls from the reimaged host time out whenever they leave the network.
billing-sync fails every export to the payments API, and no external hostname resolves on the host.
Customers outside the network cannot load the site: every external request times out while nginx reports healthy.
The migrated host cannot reach the internal API by name; every call fails before it leaves the machine.
8 incidents
Browser customers get connection refused on port 80 while the storefront's web container reports Up.
The launch API's container exits with status 0 a moment after every start, so nothing is serving.
Each time the postgres container is recreated, the orders database comes back empty and earlier rows are gone.
The finance database container exited ten minutes ago, and every finance API write has failed from that moment on.
The shop worker gets 'bad address' for the web backend on every job, so nothing it processes completes.
The release-api image build fails at a COPY step, so there is no image to run and nothing answers on the release port 3001.
Requests to the frontend on host port 8089 connect and then get an empty reply, although the container is Up and the port is published.
Compose marks the edge web container unhealthy even though its page answers 200, and the release gate that waits on health is blocked.
11 incidents
Every api replica is in CrashLoopBackOff and the Deployment has shown 0 of 3 available for ten minutes.
The migrated webapp has shown 0 of 1 available replicas from the moment it was applied to the new cluster four minutes ago.
Both order-scoring worker pods die within seconds of every start, and the orders.score backlog keeps climbing.
Ten minutes after tonight's release, the checkout Deployment has 0 of 2 replicas available and neither pod has ever started.
Calls to the catalog Service are refused even though both catalog pods are Running and Ready.
The payment pod has been stuck before startup for ten minutes, leaving the payment Deployment with zero ready replicas.
The report Deployment has run zero workers for ten minutes; its only pod is still waiting to be scheduled.
Both gateway pods have been Running for ten minutes and neither has ever become Ready, so no gateway replica is eligible for traffic.
Both pods from the new ledger revision crash on startup, leaving one old pod to carry every ledger write.
The payments deploy job's service account gets Forbidden on every Deployment patch, so the ledger change it carries cannot roll out.
The endpoint sensor on all five fleet-east nodes has been crash-looping.
4 incidents
The dev automation account's SSH logins from the deploy runner are refused, and scheduled deploy jobs cannot reach the host.
The stabilization hotfix is staged, but deploy.service cannot be restarted to ship it.
A security scan flagged the deploy account's private SSH key as readable beyond its owner on the shared deploy host.
Login, profiles and matchmaking fail for every player through the public gateway while its health checks stay green.
8 incidents
Checkout's error rate jumped from about 0.2% to 23% within minutes of the latest configuration release.
The retry change that passed checkout testing is absent from main, so the release build would ship the old behaviour.
main is clean, the tested retry change is absent from it, and no branch anywhere contains that commit.
Merging release/config into main stops on a conflict in deploy/config.yaml, so the checkout retry change cannot ship.
A release token file is tracked at the tip of main, and the credential has to be treated as exposed until it is revoked.
main still carries upstream_timeout_seconds: 2 in its committed deploy/status.yaml.
A third of partner API calls are answered with HTTP 429 while the agreed rate-limit hotfix waits to deploy.
Since this morning's 2.4.0 deploy, invoice uploads larger than 2 MB are rejected with HTTP 413.
1 incident
The deploy workflow is rejected before any job is scheduled, so none of tonight's changes can reach production.
4 incidents
Every order API request through the edge returns 502 Bad Gateway, so checkout cannot submit orders.
The web host refuses every connection on port 80, taking the site down for visitors routed to it.
Customer sites proxied through edge-fra-07 answer 502 Bad Gateway for almost every request while the edge host's CPU sits at 100%.
Every HTTPS request to the storefront fails certificate validation, so browsers and the mobile app refuse to connect.
2 incidents
API dashboards have gone flat: Prometheus reports its api scrape target down while the exporter answers from the host.
Critical alerts reach the alert router and are accepted, but no on-call page has gone out for over an hour.
1 incident
GET, LIST and PUT requests fail for every customer in nbx-east-1, and new uploads cannot be placed.
One outage, several escalating stages
5 stages
One production incident, five escalating Linux recoveries across config, routing, service control, access, and stabilization.
4 stages
Recover a broken delivery path from container build, to runtime networking, to persistence, to cluster health.
4 stages
Work through outbound access, local name resolution, DNS repair, and firewall recovery as one incident chain.
5 stages
Bring a troubled cluster release back from image pulls, configuration errors, memory kills, empty Services, and unscheduled workers.
7 stages
Contain a committed token, then recover a bad deploy, lost Git work, a config conflict, and a broken image build.
3 stages
Restore a misleading container health signal, a missing scrape target, and a paging route that hides alerts.
4 stages
Contain an exposed key and restore only the access responders need.
Free, instant, and it works on your phone. No signup: start as a guest and save your progress later.