LLenny's Podcast
← All frameworks
CommunicationRaiza Martin (Senior Product Manager, AI @ Google Labs)

Read Before You Retract: AI Incident Triage

When your AI product goes viral for something alarming, diagnose the audience's reaction before you touch the product

Difficulty
Moderate
Time to result
~days to results
Steps
5
Confidence
89%

When NotebookLM's AI hosts appeared to 'realise' they were AI and panic on-air, the clip exploded across Reddit and Twitter. Martin's response was not to pull the feature or issue a statement, but to spend a morning reading the reaction to determine whether the public had understood what actually happened — and only then to speak publicly. The framework separates the safety question from the perception question and answers the perception question with evidence.

Origin

Raiza Martin's handling of the viral 'the AI hosts realise they're AI' Audio Overview clip in 2024, which turned out to be a user-supplied show-note instructing the hosts to act it out.

Core principles

  • 01The first question is not 'is this bad?' but 'what is the attitude of the world toward this?'
  • 02Users trying to break your product is natural human curiosity, not an attack.
  • 03A viral incident is a fork in the road: retract, ignore, or engage — and you cannot choose without data.
  • 04If the audience already understands the mechanism, transparency is safe and cheap.
  • 05Reserve retraction for genuine unsafety, and keep the tested-areas list growing with every surprise.

How to run it

  1. 1

    Consume the artefact before the commentary

    Listen to or view the actual output that went viral, unmediated by the reaction. Martin heard the audio first, before reading any comments.

    Watch out Reacting to a summary of the outrage rather than the artefact leads to over-correction.

  2. 2

    Diagnose the mechanism

    Establish what actually caused the output. In this case a user had uploaded show notes telling the hosts to act out the end of the show — the model was following the source, not becoming self-aware.

    Pro tip Source-grounded products have a strong defence available: the output traces to the input.

    Watch out If you cannot explain the mechanism, you are not yet allowed to reassure anyone.

  3. 3

    Read the room, deliberately and at length

    Spend real time in the actual threads — Reddit, X, comments — determining whether the audience understands the mechanism or believes the alarming interpretation. Martin spent most of a Saturday or Sunday morning doing exactly this.

    Pro tip The signal you are looking for is whether users are explaining the mechanism to each other in the comments. If they are, you can be transparent rather than defensive.

    Watch out Do not delegate this read to a comms summary; the nuance of whether people 'get it' is only visible in the raw threads.

  4. 4

    Choose the fork with evidence

    Decide among pull it back, say nothing, or address it publicly. Martin concluded that people got it — the response showed users attributing the output to the sources — which gave her the confidence to address it publicly on Twitter rather than retract the feature.

    Pro tip Publicly acknowledging that you have seen it and understand it ('I've seen it and I get it') converts an incident into a trust deposit.

    Watch out Retraction is reserved for outputs that are genuinely unsafe, not merely embarrassing or surprising.

  5. 5

    Feed the surprise into the test suite

    Whatever the outcome, add the scenario to red-team test cases. Martin's stated stance is that new situations you did not think of get added to the tests, and only genuine unsafety triggers a pull-back.

    Pro tip Treat jailbreak creativity as a free extension of your red-team coverage.

In the wild

The hosts realise they're AI

A clip of the AI hosts panicking about being AI blew up on Reddit and Twitter over a weekend. Martin listened to it first, spent the morning reading comments, and established that users understood the behaviour came from the uploaded sources — a user had instructed the hosts to act it out.

She addressed it publicly on Twitter instead of pulling the feature, and the incident amplified rather than damaged the product's reputation.

The poop-and-fart upload

A user uploaded a document consisting of the words 'poop' and 'fart' repeated at length. Martin, about to go to bed, decided she had to listen immediately in case it was a genuine problem she would have to work on that night.

The hosts produced a genuinely insightful, funny analysis; Martin was delighted rather than alarmed, and the clip became one of the product's best-loved artefacts — no intervention needed.

Common mistakes

Pulling the feature at the first alarming clip

The instinct on going viral for the wrong reason is to retract. Had Martin done that, she would have removed a working feature over a behaviour the audience already understood as source-driven.

Treating jailbreak attempts as adversarial rather than curious

Martin frames users pushing the product into weird places as a natural part of human curiosity when people experience a technology for the first time. Reading it as attack behaviour produces defensive product decisions and hostile comms.

Speaking before you know whether people 'get it'

The decision to address it publicly was contingent on evidence that the audience already understood the mechanism. Publishing an explanation into an audience that believes something else risks amplifying the alarming interpretation.

Is it for you?

Best for

PMs and founders shipping generative AI products to the public who will inevitably face a viral, alarming, out-of-distribution output

Not ideal for

Genuinely unsafe outputs (harm, abuse, exposure), where retraction and escalation must precede any perception analysis

From the transcript

I remember thinking what is the attitude of like the world now how do we feel about this type of audio it was really the…

43:30

of course they're going to try to do things that are maybe something that we didn't think about we didn't think they would do or…

44:00

that for me was enough for me to be able to say like okay I'm going to address this publicly and say something on Twitter

45:00

if there was ever like a scenario where we're like oh this feels pretty unsafe we would pull it back

45:30

From the episode

Behind the product: NotebookLM

Raiza Martin (Senior Product Manager, AI @ Google Labs)