True Causes
← lobby
Why does the same outage keep recurring?
tap any completed stage to explore it
Generate
Clarify
Cluster
Select
Structure
Map
7
Actions
walkthrough · not counted
anyone with the link · listed publicly · federation on
30 people · 30 ideas · 2 groups
Logic → · record
THE QUESTION

Why does the same outage keep recurring?

closed · back to Actions
Viewing as guest. Enter a handle to join and take part.
close ✕
Reliability has no budget line of its own
To be precise: this is about the process, not the person on call that night.
cluster 4
Nothing said yet.
Success is measured in features shipped per quarter
Post-mortems assign actions to people who were not in the room
Scope: the services this team runs, not the whole platform.
The on-call rota has the same two people on it most weeks
Runbooks are out of date within a month of being written
To be precise: this is about the process, not the person on call that night.
Feature work always outranks reliability work in planning
Scope: the services this team runs, not the whole platform.
Post-mortem actions are written up and then never scheduled
cluster 2discuss
Every team has its own logging format
cluster 4discuss
Load tests were last run before the customer base doubled
There is no time set aside after an incident, so the write-up is done at midnight
One engineer knows how the payment path actually works
Scope: the services this team runs, not the whole platform.
Dependencies are upgraded only when something breaks
Meaning the repeat outages, not the one-off hardware failure in March.
Tests that fail intermittently are retried until green
Migrations are run by hand from a laptop
cluster 3discuss
Error budgets exist on a slide and nowhere else
Alerts fire so often that on-call mutes them
I mean the pattern over the last two quarters, not a single incident.
cluster 5discuss
Customers report outages before monitoring does
I mean the pattern over the last two quarters, not a single incident.
cluster 5discuss
The staging environment does not resemble production
Rollbacks take longer than the outage they are meant to end
To be precise: this is about the process, not the person on call that night.
Deploys go out on Friday afternoons because that is when the sprint ends
Capacity is added after the outage, never before
The same three services cause most of the pages
cluster 1discuss
Leadership asks for a root cause within the hour, so the first plausible one wins
I mean the pattern over the last two quarters, not a single incident.
The incident channel fills with people asking for status instead of giving it
Reliability has no budget line of its own
To be precise: this is about the process, not the person on call that night.
Retries are unbounded, so a slow dependency becomes a flood
Feature flags are never cleaned up, so nobody knows which paths are live
Meaning the repeat outages, not the one-off hardware failure in March.
cluster 2discuss
Config lives in five places and drifts between them
I mean the pattern over the last two quarters, not a single incident.
cluster 3discuss
Nobody owns the shared database schema
To be precise: this is about the process, not the person on call that night.
Incidents are declared late because declaring one feels like blame
The service map in the wiki is two years old
cluster 6discuss