True Causes
← lobby
Why does the same outage keep recurring?
tap any completed stage to explore it
Generate
Clarify
Cluster
Select
Structure
Map
7
Actions
walkthrough · not counted
anyone with the link · listed publicly · federation on
30 people · 30 ideas · 2 groups
Logic → · record
THE QUESTION

Why does the same outage keep recurring?

Propose what to do, then judge whether each action would move each root.

30 at the table (29 simulated) · 3 actions proposed against 6 root causes. Add one, then judge whether each action moves each cause.
Viewing as guest. Enter a handle to join and take part.
Score the roots first
why
The method's own practice, before any action is proposed: each root cause is scored 1–5 on how feasible it is to address, how much impact addressing it would have, and how likely it is to resolve by itself with no intervention. Priority goes to roots with high impact, high feasibility and low likelihood of fixing themselves — priority = impact + feasibility − likelihood. Scores are averaged across the group and shown with how many scored.
1 = low, 5 = high. Your own scores; the group's averages beside them.
root causefeasibleimpactfixes itselfpriority
The service map in the wiki is two years old
Retries are unbounded, so a slow dependency becomes a flood
Feature work always outranks reliability work in planning
One engineer knows how the payment path actually works
Post-mortems assign actions to people who were not in the room
Tests that fail intermittently are retried until green
Would each action move each root cause?
Down the side: the actions people proposed. Across the top: the root causes the map found. Every square is one vote on one question — would this action significantly reduce that cause? One side has to reach 75% of the votes cast to settle it; short of that the square stays split.
why
Answered at 75% of votes cast, like every other question here. Split means people voted and neither side reached that bar — it is a result, not a queue, and it stays split until enough people change their minds. A square nobody settles stays undecided; it never turns into a no. Judging each pair on its own is what stops a popular action being credited with fixing things it does not touch.
ACTIONROOT CAUSES · does the action on the left reduce it?
The service map in the wiki is two years oldRetries are unbounded, so a slow dependency becomes a floodFeature work always outranks reliability work in planningOne engineer knows how the payment path actually worksPost-mortems assign actions to people who were not in the roomTests that fail intermittently are retried until green
Reserve one engineer-week per sprint for post-mortem actions, first in the sprint, not last discuss✗ no, it does not
4 yes / 25 no · 22 of 29 settles it
✓ yes, it reduces this
25 yes / 4 no · 22 of 29 settles it
✗ no, it does not
2 yes / 27 no · 22 of 29 settles it
✗ no, it does not
3 yes / 26 no · 22 of 29 settles it
✓ yes, it reduces this
27 yes / 2 no · 22 of 29 settles it
✗ no, it does not
7 yes / 22 no · 22 of 29 settles it
No production deploy after 14:00 on Fridays unless the on-call engineer agrees discuss✗ no, it does not
3 yes / 26 no · 22 of 29 settles it
✓ yes, it reduces this
22 yes / 7 no · 22 of 29 settles it
✓ yes, it reduces this
26 yes / 3 no · 22 of 29 settles it
⚑ the group split
21 yes / 8 no · 22 of 29 settles it
✗ no, it does not
6 yes / 23 no · 22 of 29 settles it
✗ no, it does not
3 yes / 26 no · 22 of 29 settles it
Give the shared schema one named owner and a migration tool with rollback in focus ↑ · close ✕✗ no, it does not
6 yes / 23 no · 22 of 29 settles it
✓ yes, it reduces this
24 yes / 5 no · 22 of 29 settles it
✗ no, it does not
3 yes / 26 no · 22 of 29 settles it
✓ yes, it reduces this
24 yes / 5 no · 22 of 29 settles it
✓ yes, it reduces this
24 yes / 5 no · 22 of 29 settles it
✓ yes, it reduces this
27 yes / 2 no · 22 of 29 settles it
Highest leverage so far
Give the shared schema one named owner and a migration tool with rollback reduces 4 of 6 root causes
Reserve one engineer-week per sprint for post-mortem actions, first in the sprint, not last reduces 2 of 6 root causes
No production deploy after 14:00 on Fridays unless the on-call engineer agrees reduces 2 of 6 root causes