Storefront Stuck in Maintenance Mode
Every Northstar Shops storefront page answers 503 with the maintenance banner, and checkout has taken no orders for 20 minutes.
Free simulated incident practice
A service stops responding. A job stops running. The alert tells you what the user lost, but not which change will bring it back. These Linux exercises put you inside that gap: read the symptoms, inspect the simulated host and build an explanation before changing anything.
Separate a problem affecting one process from a problem affecting the whole host. Note what still works and when the failure began. A useful first hypothesis explains both the failures and the healthy parts of the system.
Use the terminal, files and logs available in the scenario to test your hypothesis. Check the service state and the environment it runs in. Treat a familiar error message as evidence to investigate, rather than proof of a particular fix.
After a small change, repeat the checks that showed the failure. A command succeeding is not the same as the service recovering. Finish the incident checks, then explain what evidence justified your decision.
Open a brief to see the symptoms, then launch the incident. The solution is yours to find.
Every Northstar Shops storefront page answers 503 with the maintenance banner, and checkout has taken no orders for 20 minutes.
The nightly backup logs a finished archive every run, yet BackupFailure has paged the night shift three nights running.
The capacity forecast projects the rebuilt cache host's disk full within a week at the cache log's usual growth rate.
The 02:00 content sync never published, and product pages are still showing yesterday's prices.
Three operators report that ll output on the ops host differs from what the handoff runbooks show.
The internal incident notes service is down, leaving responders to coordinate the outage from chat scrollback.
Eight jobs are stuck in the build queue, none has reached the compile step in 40 minutes, and deploys are frozen.
QueueWorkerCritical has paged 42 times in the last hour, although the queue workers report zero lag.
The nightly orders database backup exits successfully but leaves nothing in /backups.
The report API does not answer, and the report dashboard shows errors ahead of Finance's 10:00 run.
Batch exports fail with 'no space left on device', and the batch host's root filesystem is at 100%.
The metrics API will not stay up after its 09:18 restart, and the dashboards that query it show no data.
Host CPU has sat above 95% for ten minutes, and storefront response times are climbing with it.
forgeline.dev answers 500 on the web UI and API, and its readiness probe fails from every region.
The process list shows one order-router process near 98% CPU while the other seven sit between 2% and 5%.
Depot scan uploads are being refused.
The Larder delivery marketplace has accepted no stock feed from Greengage.
FixOps runs simulated systems and a terminal designed for its scenarios. It does not provision a real server or cluster. Use it to rehearse investigation and recovery, and use a real environment when you need unrestricted tooling or production experience.
There is no interview pass guarantee or certification. You can start as a guest in the browser or use the Android app. If the concepts feel unfamiliar, work through a learning track before trying another incident.
Free, instant, and it works on your phone. No signup: start as a guest and save your progress later.