True Causes
← lobby
Why does the same outage keep recurring?
tap any completed stage to explore it
Generate
Clarify
Cluster
Select
Structure
Map
7
Actions
walkthrough · not counted
anyone with the link · listed publicly · federation on
30 people · 30 ideas · 2 groups
Logic → · record
THE QUESTION

Why does the same outage keep recurring?

Does one influence the other? Three quarters of the votes cast decide.

closed · back to Actions
Viewing as guest. Enter a handle to join and take part.
Questions are shared out between groups, so you get the ones your group was given — not all 180.
why
Nobody could answer several hundred pairs. Each pair goes to one group; a pair that splits is then referred to other groups, so it can appear in your list later even though it started elsewhere. The settled count includes every group's work, which is why it climbs faster than your own list shrinks.
176 / 180 settled 0 for you to answer
Last tally: 176 decided · 1 escalated to cross-group panels · 0 contested · 3 below quorum — panels need at least max(2, 75% of panel) votes cast; your smallest panel has 15 members.
close ✕
Reliability has no budget line of its own → influences? → The staging environment does not resemble production
Nothing said yet.
Completed phase — the record of pairs and votes stands above.Logic sidecar ↗
Reliability has no budget line of its ownFeature flags are never cleaned up, so nobody knows which paths are liveopendiscuss
The same three services cause most of the pagesopendiscuss
The incident channel fills with people asking for status instead of giving itThe service map in the wiki is two years oldopendiscuss
The staging environment does not resemble productionRollbacks take longer than the outage they are meant to endopen escalateddiscuss
Tests that fail intermittently are retried until greenCustomers report outages before monitoring does✓ yesdiscuss
Alerts fire so often that on-call mutes themCustomers report outages before monitoring does✓ yesdiscuss
Alerts fire so often that on-call mutes themEvery team has its own logging format✓ yesdiscuss
Every team has its own logging formatCustomers report outages before monitoring does✓ yesdiscuss
The on-call rota has the same two people on it most weeksAlerts fire so often that on-call mutes them✓ yesdiscuss
Feature work always outranks reliability work in planningPost-mortem actions are written up and then never scheduled✓ yesdiscuss
One engineer knows how the payment path actually worksThe on-call rota has the same two people on it most weeks✓ yesdiscuss
The on-call rota has the same two people on it most weeksPost-mortem actions are written up and then never scheduled✓ yesdiscuss
Post-mortems assign actions to people who were not in the roomPost-mortem actions are written up and then never scheduled✓ yesdiscuss
The service map in the wiki is two years oldThe staging environment does not resemble production✓ yesdiscuss
Rollbacks take longer than the outage they are meant to endReliability has no budget line of its own✓ yesdiscuss
Reliability has no budget line of its ownThe staging environment does not resemble production✓ yesin focus ↑ · close ✕
Retries are unbounded, so a slow dependency becomes a floodThe same three services cause most of the pages✓ yesdiscuss
Reliability has no budget line of its ownThe incident channel fills with people asking for status instead of giving it✓ yesdiscuss
The same three services cause most of the pagesThe staging environment does not resemble production✓ yesdiscuss
The service map in the wiki is two years oldFeature flags are never cleaned up, so nobody knows which paths are live✓ yesdiscuss
Feature flags are never cleaned up, so nobody knows which paths are liveRollbacks take longer than the outage they are meant to end✓ yesdiscuss
Feature flags are never cleaned up, so nobody knows which paths are liveThe same three services cause most of the pages✓ yesdiscuss
The on-call rota has the same two people on it most weeksCustomers report outages before monitoring does✓ inferreddiscuss
Retries are unbounded, so a slow dependency becomes a floodThe staging environment does not resemble production✓ inferreddiscuss
Post-mortem actions are written up and then never scheduledEvery team has its own logging format✗ nodiscuss
Every team has its own logging formatError budgets exist on a slide and nowhere else✗ nodiscuss
Feature work always outranks reliability work in planningCustomers report outages before monitoring does✗ nodiscuss
Error budgets exist on a slide and nowhere elseTests that fail intermittently are retried until green✗ nodiscuss
Error budgets exist on a slide and nowhere elseOne engineer knows how the payment path actually works✗ nodiscuss
Error budgets exist on a slide and nowhere elsePost-mortems assign actions to people who were not in the room✗ nodiscuss
Customers report outages before monitoring doesPost-mortems assign actions to people who were not in the room✗ nodiscuss
Post-mortem actions are written up and then never scheduledAlerts fire so often that on-call mutes them✗ nodiscuss
Post-mortems assign actions to people who were not in the roomAlerts fire so often that on-call mutes them✗ nodiscuss
Error budgets exist on a slide and nowhere elseCustomers report outages before monitoring does✗ nodiscuss
Customers report outages before monitoring doesThe on-call rota has the same two people on it most weeks✗ nodiscuss
Post-mortem actions are written up and then never scheduledPost-mortems assign actions to people who were not in the room✗ nodiscuss
Feature work always outranks reliability work in planningTests that fail intermittently are retried until green✗ nodiscuss
Tests that fail intermittently are retried until greenThe on-call rota has the same two people on it most weeks✗ nodiscuss
Post-mortem actions are written up and then never scheduledThe on-call rota has the same two people on it most weeks✗ nodiscuss
Feature work always outranks reliability work in planningError budgets exist on a slide and nowhere else✗ nodiscuss
Tests that fail intermittently are retried until greenEvery team has its own logging format✗ nodiscuss
The on-call rota has the same two people on it most weeksTests that fail intermittently are retried until green✗ nodiscuss
Post-mortems assign actions to people who were not in the roomEvery team has its own logging format✗ nodiscuss
Feature work always outranks reliability work in planningPost-mortems assign actions to people who were not in the room✗ nodiscuss
Tests that fail intermittently are retried until greenAlerts fire so often that on-call mutes them✗ nodiscuss
Post-mortems assign actions to people who were not in the roomError budgets exist on a slide and nowhere else✗ nodiscuss
Alerts fire so often that on-call mutes themOne engineer knows how the payment path actually works✗ nodiscuss
Customers report outages before monitoring doesOne engineer knows how the payment path actually works✗ nodiscuss
Alerts fire so often that on-call mutes themPost-mortem actions are written up and then never scheduled✗ nodiscuss
Post-mortems assign actions to people who were not in the roomThe on-call rota has the same two people on it most weeks✗ nodiscuss
One engineer knows how the payment path actually worksAlerts fire so often that on-call mutes them✗ nodiscuss
Customers report outages before monitoring doesAlerts fire so often that on-call mutes them✗ nodiscuss
Error budgets exist on a slide and nowhere elseThe on-call rota has the same two people on it most weeks✗ nodiscuss
Feature work always outranks reliability work in planningEvery team has its own logging format✗ nodiscuss
Customers report outages before monitoring doesPost-mortem actions are written up and then never scheduled✗ nodiscuss
Customers report outages before monitoring doesEvery team has its own logging format✗ nodiscuss
Customers report outages before monitoring doesError budgets exist on a slide and nowhere else✗ nodiscuss
Post-mortem actions are written up and then never scheduledOne engineer knows how the payment path actually works✗ nodiscuss
Customers report outages before monitoring doesTests that fail intermittently are retried until green✗ nodiscuss
Every team has its own logging formatThe on-call rota has the same two people on it most weeks✗ nodiscuss
Every team has its own logging formatPost-mortems assign actions to people who were not in the room✗ nodiscuss
Every team has its own logging formatPost-mortem actions are written up and then never scheduled✗ nodiscuss
Post-mortem actions are written up and then never scheduledTests that fail intermittently are retried until green✗ nodiscuss
The on-call rota has the same two people on it most weeksError budgets exist on a slide and nowhere else✗ nodiscuss
Alerts fire so often that on-call mutes themThe on-call rota has the same two people on it most weeks✗ nodiscuss
One engineer knows how the payment path actually worksTests that fail intermittently are retried until green✗ nodiscuss
Tests that fail intermittently are retried until greenPost-mortems assign actions to people who were not in the room✗ nodiscuss
Feature work always outranks reliability work in planningAlerts fire so often that on-call mutes them✗ nodiscuss
Customers report outages before monitoring doesFeature work always outranks reliability work in planning✗ nodiscuss
One engineer knows how the payment path actually worksCustomers report outages before monitoring does✗ nodiscuss
Post-mortems assign actions to people who were not in the roomCustomers report outages before monitoring does✗ nodiscuss
Error budgets exist on a slide and nowhere elseFeature work always outranks reliability work in planning✗ nodiscuss
Alerts fire so often that on-call mutes themPost-mortems assign actions to people who were not in the room✗ nodiscuss
One engineer knows how the payment path actually worksPost-mortem actions are written up and then never scheduled✗ nodiscuss
The on-call rota has the same two people on it most weeksOne engineer knows how the payment path actually works✗ nodiscuss
Tests that fail intermittently are retried until greenOne engineer knows how the payment path actually works✗ nodiscuss
The on-call rota has the same two people on it most weeksEvery team has its own logging format✗ nodiscuss
Alerts fire so often that on-call mutes themTests that fail intermittently are retried until green✗ nodiscuss
Every team has its own logging formatOne engineer knows how the payment path actually works✗ nodiscuss
The on-call rota has the same two people on it most weeksPost-mortems assign actions to people who were not in the room✗ nodiscuss
Post-mortems assign actions to people who were not in the roomTests that fail intermittently are retried until green✗ nodiscuss
Every team has its own logging formatAlerts fire so often that on-call mutes them✗ nodiscuss
The on-call rota has the same two people on it most weeksFeature work always outranks reliability work in planning✗ nodiscuss
Tests that fail intermittently are retried until greenFeature work always outranks reliability work in planning✗ nodiscuss
Tests that fail intermittently are retried until greenPost-mortem actions are written up and then never scheduled✗ nodiscuss
Error budgets exist on a slide and nowhere elseAlerts fire so often that on-call mutes them✗ nodiscuss
Error budgets exist on a slide and nowhere elsePost-mortem actions are written up and then never scheduled✗ nodiscuss
Post-mortems assign actions to people who were not in the roomFeature work always outranks reliability work in planning✗ nodiscuss
One engineer knows how the payment path actually worksFeature work always outranks reliability work in planning✗ nodiscuss
Error budgets exist on a slide and nowhere elseEvery team has its own logging format✗ nodiscuss
Tests that fail intermittently are retried until greenError budgets exist on a slide and nowhere else✗ nodiscuss
Post-mortem actions are written up and then never scheduledError budgets exist on a slide and nowhere else✗ nodiscuss
Post-mortem actions are written up and then never scheduledCustomers report outages before monitoring does✗ nodiscuss
Every team has its own logging formatTests that fail intermittently are retried until green✗ nodiscuss
One engineer knows how the payment path actually worksPost-mortems assign actions to people who were not in the room✗ nodiscuss
Feature work always outranks reliability work in planningOne engineer knows how the payment path actually works✗ nodiscuss
Every team has its own logging formatFeature work always outranks reliability work in planning✗ nodiscuss
One engineer knows how the payment path actually worksError budgets exist on a slide and nowhere else✗ nodiscuss
Alerts fire so often that on-call mutes themError budgets exist on a slide and nowhere else✗ nodiscuss
Post-mortems assign actions to people who were not in the roomOne engineer knows how the payment path actually works✗ nodiscuss
Alerts fire so often that on-call mutes themFeature work always outranks reliability work in planning✗ nodiscuss
One engineer knows how the payment path actually worksEvery team has its own logging format✗ nodiscuss
Feature work always outranks reliability work in planningThe on-call rota has the same two people on it most weeks✗ nodiscuss
Post-mortem actions are written up and then never scheduledFeature work always outranks reliability work in planning✗ nodiscuss
The same three services cause most of the pagesThe service map in the wiki is two years old✗ nodiscuss
Rollbacks take longer than the outage they are meant to endThe same three services cause most of the pages✗ nodiscuss
Config lives in five places and drifts between themRollbacks take longer than the outage they are meant to end✗ nodiscuss
The staging environment does not resemble production✗ nodiscuss
The incident channel fills with people asking for status instead of giving itConfig lives in five places and drifts between them✗ nodiscuss
Reliability has no budget line of its ownRetries are unbounded, so a slow dependency becomes a flood✗ nodiscuss
The incident channel fills with people asking for status instead of giving itRollbacks take longer than the outage they are meant to end✗ nodiscuss
The same three services cause most of the pagesThe incident channel fills with people asking for status instead of giving it✗ nodiscuss
The service map in the wiki is two years oldRetries are unbounded, so a slow dependency becomes a flood✗ nodiscuss
Config lives in five places and drifts between themFeature flags are never cleaned up, so nobody knows which paths are live✗ nodiscuss
The incident channel fills with people asking for status instead of giving itFeature flags are never cleaned up, so nobody knows which paths are live✗ nodiscuss
Retries are unbounded, so a slow dependency becomes a floodThe service map in the wiki is two years old✗ nodiscuss
Reliability has no budget line of its own✗ nodiscuss
The service map in the wiki is two years oldRollbacks take longer than the outage they are meant to end✗ nodiscuss
Feature flags are never cleaned up, so nobody knows which paths are liveReliability has no budget line of its own✗ nodiscuss
Feature flags are never cleaned up, so nobody knows which paths are liveThe incident channel fills with people asking for status instead of giving it✗ nodiscuss
Rollbacks take longer than the outage they are meant to end✗ nodiscuss
The incident channel fills with people asking for status instead of giving itReliability has no budget line of its own✗ nodiscuss
Config lives in five places and drifts between themThe staging environment does not resemble production✗ nodiscuss
Config lives in five places and drifts between them✗ nodiscuss
Feature flags are never cleaned up, so nobody knows which paths are live✗ nodiscuss
Reliability has no budget line of its ownConfig lives in five places and drifts between them✗ nodiscuss
The incident channel fills with people asking for status instead of giving itRetries are unbounded, so a slow dependency becomes a flood✗ nodiscuss
Rollbacks take longer than the outage they are meant to endThe incident channel fills with people asking for status instead of giving it✗ nodiscuss
The same three services cause most of the pagesReliability has no budget line of its own✗ nodiscuss
Feature flags are never cleaned up, so nobody knows which paths are live✗ nodiscuss
Feature flags are never cleaned up, so nobody knows which paths are liveThe service map in the wiki is two years old✗ nodiscuss
Config lives in five places and drifts between themReliability has no budget line of its own✗ nodiscuss
The staging environment does not resemble productionFeature flags are never cleaned up, so nobody knows which paths are live✗ nodiscuss
Rollbacks take longer than the outage they are meant to endFeature flags are never cleaned up, so nobody knows which paths are live✗ nodiscuss
Retries are unbounded, so a slow dependency becomes a floodConfig lives in five places and drifts between them✗ nodiscuss
The service map in the wiki is two years oldReliability has no budget line of its own✗ nodiscuss
The service map in the wiki is two years oldConfig lives in five places and drifts between them✗ nodiscuss
The staging environment does not resemble productionConfig lives in five places and drifts between them✗ nodiscuss
Rollbacks take longer than the outage they are meant to endThe service map in the wiki is two years old✗ nodiscuss
The staging environment does not resemble productionThe service map in the wiki is two years old✗ nodiscuss
The staging environment does not resemble productionRetries are unbounded, so a slow dependency becomes a flood✗ nodiscuss
The same three services cause most of the pages✗ nodiscuss
Retries are unbounded, so a slow dependency becomes a floodFeature flags are never cleaned up, so nobody knows which paths are live✗ nodiscuss
The service map in the wiki is two years old✗ nodiscuss
Config lives in five places and drifts between themRetries are unbounded, so a slow dependency becomes a flood✗ nodiscuss
The staging environment does not resemble production✗ nodiscuss
Reliability has no budget line of its own✗ nodiscuss
The same three services cause most of the pagesConfig lives in five places and drifts between them✗ nodiscuss
The incident channel fills with people asking for status instead of giving it✗ nodiscuss
Feature flags are never cleaned up, so nobody knows which paths are liveThe staging environment does not resemble production✗ nodiscuss
Retries are unbounded, so a slow dependency becomes a floodThe incident channel fills with people asking for status instead of giving it✗ nodiscuss
Config lives in five places and drifts between them✗ nodiscuss
The staging environment does not resemble productionThe same three services cause most of the pages✗ nodiscuss
Reliability has no budget line of its ownRollbacks take longer than the outage they are meant to end✗ nodiscuss
Config lives in five places and drifts between themThe same three services cause most of the pages✗ nodiscuss
The staging environment does not resemble productionThe incident channel fills with people asking for status instead of giving it✗ nodiscuss
Rollbacks take longer than the outage they are meant to endThe staging environment does not resemble production✗ nodiscuss
Rollbacks take longer than the outage they are meant to end✗ nodiscuss
Config lives in five places and drifts between themThe service map in the wiki is two years old✗ nodiscuss
The same three services cause most of the pagesRetries are unbounded, so a slow dependency becomes a flood✗ nodiscuss
Rollbacks take longer than the outage they are meant to endConfig lives in five places and drifts between them✗ nodiscuss
Reliability has no budget line of its ownThe service map in the wiki is two years old✗ nodiscuss
Retries are unbounded, so a slow dependency becomes a flood✗ nodiscuss
Feature flags are never cleaned up, so nobody knows which paths are liveConfig lives in five places and drifts between them✗ nodiscuss
Config lives in five places and drifts between themThe incident channel fills with people asking for status instead of giving it✗ nodiscuss
Feature flags are never cleaned up, so nobody knows which paths are liveRetries are unbounded, so a slow dependency becomes a flood✗ nodiscuss
The service map in the wiki is two years old✗ nodiscuss
The incident channel fills with people asking for status instead of giving itThe same three services cause most of the pages✗ nodiscuss
Rollbacks take longer than the outage they are meant to endRetries are unbounded, so a slow dependency becomes a flood✗ nodiscuss
Reliability has no budget line of its ownThe same three services cause most of the pages✗ nodiscuss
Retries are unbounded, so a slow dependency becomes a floodRollbacks take longer than the outage they are meant to end✗ nodiscuss
The same three services cause most of the pagesFeature flags are never cleaned up, so nobody knows which paths are live✗ nodiscuss
The service map in the wiki is two years oldThe incident channel fills with people asking for status instead of giving it✗ nodiscuss
The service map in the wiki is two years oldThe same three services cause most of the pages✗ nodiscuss
The incident channel fills with people asking for status instead of giving it✗ nodiscuss
The staging environment does not resemble productionReliability has no budget line of its own✗ nodiscuss
The incident channel fills with people asking for status instead of giving itThe staging environment does not resemble production✗ nodiscuss
The same three services cause most of the pagesRollbacks take longer than the outage they are meant to end✗ nodiscuss
Retries are unbounded, so a slow dependency becomes a flood✗ nodiscuss
Retries are unbounded, so a slow dependency becomes a floodReliability has no budget line of its own✗ nodiscuss