Briefing
forgeline.dev started answering 500 on every page at 09:07, about twenty-five minutes after the database team began rebuilding the standby that fell behind during this morning's spam wave.
The main PostgreSQL cluster on pg-core-01 will not start, writes are frozen at the load balancer, and git pushes over SSH still work only.
Dev wants to reuse the standby's half-finished copy.
Inspired by a real outage: Postmortem of database outage of January 31 (GitLab, 2017). Names and details are fictionalized.
The fix is yours to find: hints, the recovery checks and the debrief unlock inside the incident.