Fail-State Goaling
Name your worst-case experiences, set a team's goal on eradicating them, and hunt for the data you don't have
- Difficulty
- Moderate
- Time to result
- ~months to results
- Steps
- 5
- Confidence
- 94%
Averages hide catastrophes. DoorDash defines its disaster outcomes explicitly (a 'never delivered' order), gives a cross-functional team the literal goal of eradicating them, and quantifies their true cost — refund, repurchase, second courier, and all the future orders a churned customer never places. It pairs this with a habit of asking what data is missing entirely, since some failures (like a login error) remove the user from the denominator before they ever appear in a dataset.
Origin
Jessica Lachs, DoorDash quality analytics — home of the deliberately literal metric name 'never delivered'.
Core principles
- 01Rarity is not unimportance — a fail state can cost far more than its frequency suggests
- 02Averages and percentile quality metrics structurally hide the tail
- 03Churn hides the biggest cost: the lost future orders are never observed
- 04Some failures never enter the dataset at all — ask what data you don't have
- 05Name metrics with brutal literalness so everyone knows what they mean
How to run it
- 1
Enumerate your fail states
List the experiences that go terribly wrong, not just the ones that go slightly below average: the order never delivered, the user who cannot log in, the catastrophic support failure.
Pro tip Name them plainly. 'Never delivered' beats any acronym — clarity above all else.
Watch out If you only monitor averages and delivery-time percentiles, these events will literally never surface.
- 2
Quantify their true cost
Cost out the full impact: refund, repurchase, second delivery, plus the customer's entire stream of subsequent orders lost to churn. This is the number that makes a rare event worth a team.
Pro tip The churn cost is invisible by construction — model it, because the lost orders will never show up as data.
- 3
Give a cross-functional team the eradication goal
Stand up a team spanning analytics, product, engineering, and ops whose stated goal is to eradicate the fail state — not reduce it as a side-effect of quality work.
Pro tip Accept you will never reach zero; aim to push it from a fraction of a percent to a fraction of a fraction.
- 4
Diagnose root causes and intervene mid-flight
Understand why each fail state happens — human error, fraud, systemic gaps — then build both prevention and in-flight recovery so an order that is going wrong can be rescued before it becomes a never-delivered.
- 5
Ask what data you don't have
Systematically ask which failures remove the user from the dataset. Login failures mean the user never enters the denominator, so the metric looks fine while people are being silently locked out.
Pro tip Make 'what are we missing from the denominator?' a standing question in metric reviews.
Watch out The data will not tell you it is incomplete. This has to be a deliberate habit.
In the wild
DoorDash's most extreme failure is an order that is simply never delivered. It is very rare and invisible in average delivery-time and lateness metrics — but it causes churn and costs a refund plus repurchased food plus a second Dasher. Quality analytics, product, engineering, and ops share one goal: eradicate never delivered.
→ Never-delivered rates driven down to a fraction of a fraction of a percent, protecting both consumer experience and margin far beyond what the event's frequency implies.
Users who cannot log in place no orders and generate no in-app events, so they are absent from the denominator of every engagement metric. The failure is invisible unless someone deliberately asks what data is missing.
→ Lachs treats 'what data don't we have' as a core discipline for data teams, surfacing opportunities that never appear in existing dashboards.
Common mistakes
Dismissing rare events as immaterial
Frequency is a bad proxy for cost. A rare fail state can be enormously expensive per occurrence and cause churn whose full cost is never observed in the data.
Measuring quality only on averages
Average delivery times and average consumer experience are structurally incapable of surfacing the tail. The disasters simply never show up.
Assuming the dataset is complete
Failures that block entry (login errors, signup breaks) remove people from the denominator, so metrics look healthy while users are locked out.
Is it for you?
Best for
Data, quality, and ops leaders at operational or marketplace businesses where rare catastrophic experiences drive churn and cost disproportionately
Not ideal for
Early products still fighting for baseline usage, where the average experience is itself broken and tail optimisation is premature
From the transcript
“making sure that you're looking at the edge cases and your fail States is also really important and so we often will set goals actually…”
“we have this concept of never delivered”
“their goal is eradicate never delivered”
“you're losing all of that consumer's subsequent orders and that is not necessarily observed”
“what data don't we have what data might we be missing”
From the episode
Building a world-class data org
Jessica Lachs (VP of Analytics and Data Science at DoorDash)