True Causes
← lobby
Why does the same outage keep recurring?
tap any completed stage to explore it
Generate
Clarify
Cluster
Select
Structure
Map
7
Actions
walkthrough · not counted
anyone with the link · listed publicly · federation on
30 people · 30 ideas · 2 groups
Logic → · record
THE QUESTION

Why does the same outage keep recurring?

The causes, deepest first. Disagreements stay on the map.

closed · back to Actions
Viewing as guest. Enter a handle to join and take part.
The map
Left to right: deepest causes first; outlined green boxes are the roots. Only direct links are drawn — a link implied by a path is on the record, not in the picture. Dashed red is contested. Hover for full text.
The same three services cause most of the pages → The staging environment does not resemble productionFeature work always outranks reliability work in planning → Post-mortem actions are written up and then never scheduledRollbacks take longer than the outage they are meant to end → Reliability has no budget line of its ownEvery team has its own logging format → Customers report outages before monitoring doesReliability has no budget line of its own → The incident channel fills with people asking for status instead of giReliability has no budget line of its own → The staging environment does not resemble productionPost-mortems assign actions to people who were not in the room → Post-mortem actions are written up and then never scheduledTests that fail intermittently are retried until green → Customers report outages before monitoring doesRetries are unbounded, so a slow dependency becomes a flood → The same three services cause most of the pagesAlerts fire so often that on-call mutes them → Every team has its own logging formatThe on-call rota has the same two people on it most weeks → Post-mortem actions are written up and then never scheduledThe on-call rota has the same two people on it most weeks → Alerts fire so often that on-call mutes themFeature flags are never cleaned up, so nobody knows which paths are li → The same three services cause most of the pagesFeature flags are never cleaned up, so nobody knows which paths are li → Rollbacks take longer than the outage they are meant to endOne engineer knows how the payment path actually works → The on-call rota has the same two people on it most weeksThe service map in the wiki is two years old → Feature flags are never cleaned up, so nobody knows which paths are liFeature work always outranks reliability work in planningFeature work alwaysoutranks reliability work…Post-mortems assign actions to people who were not in the roomPost-mortems assign actionsto people who were not in…Tests that fail intermittently are retried until greenTests that failintermittently are retried…Retries are unbounded, so a slow dependency becomes a floodRetries are unbounded, so aslow dependency becomes a…One engineer knows how the payment path actually worksOne engineer knows how thepayment path actually worksThe service map in the wiki is two years oldThe service map in the wikiclusterCLUSTER 6Feature flags are never cleaned up, so nobody knows which paths are liFeature flags are neverclusterCLUSTER 2The on-call rota has the same two people on it most weeksThe on-call rota has theclusterCLUSTER 1Rollbacks take longer than the outage they are meant to endRollbacks take longer thanthe outage they are meant…Alerts fire so often that on-call mutes themAlerts fire so often thatclusterCLUSTER 5The same three services cause most of the pagesThe same three servicesclusterCLUSTER 1Post-mortem actions are written up and then never scheduledPost-mortem actions areclusterCLUSTER 2Reliability has no budget line of its ownReliability has no budgetclusterCLUSTER 4Every team has its own logging formatEvery team has its ownclusterCLUSTER 4The incident channel fills with people asking for status instead of giThe incident channel fillswith people asking for…The staging environment does not resemble productionThe staging environmentdoes not resemble…Customers report outages before monitoring doesCustomers report outagesclusterCLUSTER 5
How this dialogue measures
why
The indicators the method's own practitioners use to judge a dialogue's quality (Laouris & Metcalf 2024, Table 7), computed from this record so it can be compared with the hundreds of face-to-face and virtual dialogues behind their baselines. Spreadthink is the share of ideas that got at least one vote: 100% means everyone voted only for their own, and it usually settles to 25–45% once meaning has been shared. Situational complexity falls as more ideas are structured.
ideas (N)
30 usual: 50–100
with 1+ vote / 2+ votes
28 / 24
spreadthink
92.0% usual: 25–55
demosensus
8.0%
clusters (dimensions)
6
cards in a cluster
40.0%
structured / levels / links
20 / 5 / 20
situational complexity
1.11
minds moved
cards hidden by the steward
0
pairs contested
0
The merged map
why
Every group built its own map from its own wall; these are folded into one. The same claim raised in two groups is one item, and each connection shows how many groups found it. Where groups found opposite things, that is published as a conflict.
ROOT CAUSES
d13e1db
afe8f84
5f15967
3a44a54
3201ca5
03944a4
Connections several groups found (20 of 20)
Alerts fire so often that on-call mutes themCustomers report outages before monitoring does1 group
The service map in the wiki is two years oldFeature flags are never cleaned up, so nobody knows which paths are li1 group
Alerts fire so often that on-call mutes themEvery team has its own logging format1 group
Reliability has no budget line of its ownThe staging environment does not resemble production1 group
Post-mortems assign actions to people who were not in the roomPost-mortem actions are written up and then never scheduled1 group
Feature flags are never cleaned up, so nobody knows which paths are liThe same three services cause most of the pages1 group
Feature work always outranks reliability work in planningPost-mortem actions are written up and then never scheduled1 group
The on-call rota has the same two people on it most weeksCustomers report outages before monitoring does1 group
Reliability has no budget line of its ownThe incident channel fills with people asking for status instead of gi1 group
Rollbacks take longer than the outage they are meant to endReliability has no budget line of its own1 group
Feature flags are never cleaned up, so nobody knows which paths are liRollbacks take longer than the outage they are meant to end1 group
Retries are unbounded, so a slow dependency becomes a floodThe same three services cause most of the pages1 group
Retries are unbounded, so a slow dependency becomes a floodThe staging environment does not resemble production1 group
Every team has its own logging formatCustomers report outages before monitoring does1 group
The on-call rota has the same two people on it most weeksPost-mortem actions are written up and then never scheduled1 group
One engineer knows how the payment path actually worksThe on-call rota has the same two people on it most weeks1 group
The service map in the wiki is two years oldThe staging environment does not resemble production1 group
The on-call rota has the same two people on it most weeksAlerts fire so often that on-call mutes them1 group
The same three services cause most of the pagesThe staging environment does not resemble production1 group
Tests that fail intermittently are retried until greenCustomers report outages before monitoring does1 group