LLenny's Podcast
← All frameworks
LeadershipEthan Evans (Amazon)

Buying Life One Hour at a Time (Public Failure Recovery)

Own the failure, buy trust back one hour at a time, then get face-to-face.

Difficulty
Advanced
Time to result
~days to results
Steps
6
Confidence
95%

A protocol for the hours and weeks after you visibly fail in front of powerful people. Rather than defending, disappearing, or drowning in shame, you take unqualified ownership, then re-establish trust with proactive, timeboxed updates that pre-empt micromanagement — 'buying life one hour at a time'. Once the fire is out, you deliberately move the relationship out of email and into a face-to-face encounter, because flaming someone in person is far harder than flaming them by email.

Origin

Ethan Evans' account of the Amazon Appstore launch, where the 'test drive' feature Jeff Bezos had personally highlighted in his customer launch letter failed at 6am on launch day due to a database scaling defect. Bezos, Jeff Wilke (CEO of retail), and successive layers of leadership piled onto an increasingly angry email thread. Evans survived and was promoted to VP two years later.

Core principles

  • 01Own it completely and immediately — no deflection, no 'the system failed'.
  • 02The absence of information is what triggers micromanagement; a credible update cadence is what holds it off.
  • 03Trust is not restored in one move; it is rented back in hourly increments.
  • 04Commit to the next update time, not the fix time — you can always keep an update promise.
  • 05Anger is easy at distance and hard in person; engineer proximity.
  • 06Do not chicken out of the meeting with the person who is angry at you. 'If I can't face the CEO I'd better pack my desk.'
  • 07Shame is the real career killer — most people quit or shrink long before they are actually dead in the water.
  • 08The recovery is only complete when you can name and demonstrate the lesson learned.

How to run it

  1. 1

    Own it, flatly and first

    Before explaining anything, state that it's broken, it's your fault, and you will deal with it. No hedging, no distributing blame down the chain to the engineer who wrote the code.

    Pro tip Ownership includes doing the unglamorous recovery work — sending your team home in sleep shifts, working the weekend, pulling in every favor. 'It's on you to pull out the stops even if it's uncomfortable.'

    Watch out Explaining before owning reads as excuse-making and accelerates the pile-on.

  2. 2

    Install a proactive hourly update cadence

    Send updates on a fixed clock: here is exactly where we are, here is what we will do in the next hour, and here is when you will get your next update. Get agreement on the plan, and then keep the update promise even when there is no progress.

    Pro tip Each kept update buys another hour of autonomy. The goal is to stop three or four levels of management from descending to 'help'.

    Watch out Never miss a promised update. The cadence, not the content, is what is rebuilding trust.

  3. 3

    Accept the help that arrives

    Senior people who have been on the wrong end of the same scrutiny will back-channel offers of help. Take them. In Evans' case, AWS principal engineers arrived, diagnosed a fundamentally flawed database design, and simply threw ~500 machines at it so the bad design would run anyway.

    Pro tip The immediate fix and the correct fix are different projects. Buy your way out of the outage now, fix the architecture after.

    Watch out Pride here is fatal — refusing help to protect your reputation costs you the reputation.

  4. 4

    Force a face-to-face encounter

    Once the emergency passes, deliberately put yourself physically in front of the angry executive rather than hiding. Evans arrived early to a meeting where he was a peripheral participant and sat in the chair next to where Bezos always sat.

    Pro tip Do not open with a defense. Let them speak first — given the chance to be kind in person, most people take it.

    Watch out Do not assume the same play works on everyone. Evans ran the identical playbook on Jeff Wilke and was told he had come within a hair of being fired.

  5. 5

    Name the lesson and change the behavior for years

    Forgiveness is not restored trust. Evans identified his actual defect — being an 'operational cowboy' who prioritized speed over certainty — and installed a new rule for himself and his team: 'Fear the New York Times headline.' Ask whether what you are about to do could produce a headline; if so, slow down.

    Pro tip He also killed surprise launches in favor of beta testing: 'the biggest thing I learned with surprise launches is that you're surprised by what doesn't work' — a leak from an NDA'd beta tester is a better outcome than shipping something broken.

    Watch out Do not overcorrect into paralysis. The rule is 'be really careful when a headline is possible', not 'never gamble'.

  6. 6

    Check on the person further down who actually made the error

    Find the junior person whose code or decision broke, and explicitly tell them the system failed them, not the other way round. They watched you take a beating and have concluded it was their fault.

    Pro tip Do this within days, not weeks — the shame calcifies fast.

    Watch out Evans' single regret: the new-grad engineer who wrote the unscalable code was never reassured, and left the company. 'He felt undue responsibility, and that I really regret.'

In the wild

The Appstore launch that failed in front of Bezos

The Amazon Appstore's 'test drive' feature — the exact feature Bezos had chosen to highlight in his personal customer launch letter — was still broken at 6am on launch day. Bezos emailed at 6:15am asking where the letter was; the ensuing thread pulled in Evans' boss, his boss's boss, and Jeff Wilke, and grew angrier by the hour. Evans took ownership, committed to hourly updates with a plan for each next hour, accepted AWS principal engineers who papered over the flawed database design with ~500 machines, and the following week deliberately sat next to Bezos' chair at a meeting he could have skipped.

Bezos turned at the end of the meeting and said 'so how are you doing, I bet it's been a hard week' — a human conversation that signaled forgiveness, though not restored trust. Evans re-earned the trust over the following two years and was promoted to vice president.

The Wilke meeting that revealed how close he came

Emboldened by the Bezos meeting, Evans scheduled a face-to-face with Jeff Wilke expecting the same playbook to work. Wilke instead asked whether he had known he was gambling when he launched. Evans said yes — he had judged hitting the media commitment date more important than perfect certainty.

Wilke told him he had been wrong to prioritize the date over Amazon's public reputation — but that at least he had known he was gambling: 'if you hadn't known you were gambling we'd be discussing your departure.' Evans chose stubbornness over shame and moved forward.

Common mistakes

Going quiet while you fix it

Silence is what triggers the executive to assume nobody is on it and to start micromanaging — which pulls three or four levels of management into the incident and makes it worse.

Avoiding the angry person

Skipping the meeting feels safe and is fatal. Email is where anger lives; in person, the same executive will usually choose to say something kind. If you cannot face them, you have already resigned.

Letting shame do the firing for them

Evans notes many people in his position would have quit. 'A lot of people feel they're more dead in the water than they are, because everybody makes mistakes.' The stubborn refusal to live in shame is what kept his career alive.

Forgetting the junior person who wrote the bug

Nobody yelled at the new-grad engineer, but he watched his leaders take a beating over his code and quietly left the company. Ownership means absorbing the blame publicly and then actively telling that person the system, not they, failed.

Is it for you?

Best for

Engineering and product leaders in the middle of a visible, high-stakes failure with senior executives watching, who need to stabilize both the incident and their standing.

Not ideal for

Low-stakes mistakes where hourly executive updates would be theatrical, or genuinely unethical failures where recovery is not the appropriate goal.

From the transcript

so the first thing I did was I owned it

44:00

this is exactly where we are this is what we're going to do in the next hour and this is when you'll get your next…

44:30

so I was buying life one hour at a time

45:00

when I had that thought I realized you know if I can't face the CEO I'd better pack my desk

48:00

he said second though at least you knew you were gambling if you hadn't known you were gambling we'd be discussing your departure

55:00

the biggest thing I learned with surprise launches is that you're surprised by what doesn't work

58:00

From the episode

Taking control of your career

Ethan Evans (Amazon)