LLenny's Podcast
← All episodes
Keith Coleman (VP of Product) and Jay Baxter (ML Lead)27 February 2025

An inside look at X’s Community Notes

6Frameworks
15Insights

Frameworks in this episode

Insights & moments

The myth-busts, hot takes, explainers, and tools worth keeping.

Myth Buster· 2

Myth Buster08:00

You Don't Need Fact-Checker Labels to Bootstrap a Trust System

The prevailing assumption among ML engineers was that a system like this had to be closed-source and trained on ground-truth labels from professional fact-checkers, or manipulators would overrun it. Community Notes proved you can bootstrap the whole thing with no external labels using a bridging-based agreement algorithm.

  • Conventional ML wisdom in 2020 said the system had to be closed-source and label-dependent
  • Fear was constant manipulation without ground-truth fact-checker labels
  • The bridging-based agreement approach works without any external labels
  • Cross-disagreement agreement also provides strong anti-manipulation properties

I think a room of ml Engineers would say oh you have to keep it closed Source you know people are going to be manipulating…

08:00
#machine-learning#anti-manipulation#algorithm#open-source
Myth Buster1:26:30

The Counterintuitive Win: Anonymous Contributors Are More Honest

The team originally assumed notes should carry real names for trust — and were completely wrong. The pilot's biggest lesson was that pseudonymity works better: people were afraid to be harassed writing notes under their real name, and studies plus their own data show people cross partisan boundaries far more readily when anonymous. Their many quality mechanisms kept quality high regardless.

  • The first prototypes assumed real names would build trust; that assumption was wrong
  • Under real names, people feared harassment and self-censored on controversial topics
  • People are far more willing to agree with the other side when anonymous or pseudonymous
  • Existing quality mechanisms meant pseudonymity didn't lower note quality

one people were hesitant to write a note on a controversial topic because they didn't want to get like attacked or harassed online and so…

1:27:00

it allows it allows freedom for honesty which is pretty great

1:28:00
#anonymity#psychology#product-design#honesty

Hot Take· 3

Hot Take37:30

The Myth That Managing More People Means More Impact

Keith argues the belief that a bigger org or wider scope equals more impact is a myth that traps ambitious leaders. Had he stayed running a large consumer PM team he'd have produced more OKR pages; instead a tiny team built the industry standard for handling information quality at internet scale.

  • Bigger scope / more direct reports does not equal more impact
  • Staying on the management track would have produced documents, not a world-changing product
  • Community Notes became the industry standard for dealing with misinformation
  • The classic career path rewards headcount, but working on what you love matters more

myth that that can get in people's way is the idea that the the more people you manage or something or the larger your scope…

Keith Coleman · 38:30
#leadership#career#impact#management
Hot Take1:03:30

On a Small Team, Deleting Code Matters More Than Writing It

Jay argues that a forced-small team is an advantage because it makes you delete code. Engineers tend to add little incremental wins from short A/B tests without appreciating the long-term maintenance burden they create. Auditing systems and cutting features whose upkeep cost exceeds their gains keeps a small team fast.

  • Deleting code is often more important than writing it
  • Promotion incentives push engineers to add small, ship-a-win features
  • A one-month A/B test hides the eternal maintenance cost of what you add
  • A small team is forced to audit systems and delete low-value complexity

deleting code is more important than writing it a lot of the time uh so I I think so often maybe due to promotion incentives…

Jay Baxter · 1:04:00
#engineering#small-teams#maintenance#code
Hot Take1:11:00

The Craziest Principle: There Is No Button to Take a Note Down

Keith says the foundational, most unsettling principle is that Community Notes is the voice of the people, not the company. There is deliberately no button for the company to change a note's status: if a note is rated helpful, it shows. If a note is bad enough to want to override, that's treated as a signal to fix the system, not to reach in and pull the note.

  • The product represents the voice of the people, not the company's voice
  • There is intentionally no button to change a note's status
  • If a helpful-rated note shows, the company can't take it down
  • A note bad enough to want removed is a system problem to redesign, not to override

probably the craziest one is just that this thing is going to be the voice of the people it's going to represent the voice of…

Keith Coleman · 11:30

it has to work that well if it doesn't work well enough to do that then it doesn't work

Keith Coleman · 12:00
#principles#trust#product-design#governance

Explainer· 3

Explainer06:00

How Community Notes Uses "Surprising Agreement" to Stay Neutral

Community Notes doesn't take a majority-rules vote. It looks for agreement between people who have historically disagreed with each other, and only shows a note when those opposed groups converge on it. That surprising cross-partisan agreement is what makes a note read as neutral, accurate, and well written.

  • Anyone on X can flag a misleading post and propose a note; others rate it
  • A note only shows if people who normally disagree find it helpful
  • Majority-rules voting or instant publishing would produce biased or inaccurate notes
  • Cross-partisan agreement acts as a proxy for accuracy and neutrality

if the note is found helpful by people who normally disagree with each other indicating that it's probably accurate it's probably really neutrally worded it's…

06:30

we actually look for agreement from people who have disagreed in the past uh and and what we see is when people actually have that…

07:30
#community-notes#algorithm#misinformation#trust
Explainer19:30

Only ~8% of Proposed Notes Ever Show — Because a Bad Note Is the Worst Outcome

Community Notes deliberately shows only about 7-11% of proposed notes. People ask why they don't show more, but the team holds an intentionally conservative bar because a single bad note undermines the trust the whole product depends on. Writing notes that opposing groups rate unhelpful can cost you your ability to write.

  • Roughly 8% of proposed notes get shown (range ~7-11% over time)
  • The threshold is set conservatively on purpose to keep note quality high
  • Showing a bad note is viewed as the worst possible mistake because it erodes trust
  • Writers whose notes are found unhelpful by cross-partisan raters lose and must re-earn writing ability

we probably show about 8% of notes that get proposed um I think that's it's been between it's say 7% and 10% or 11% something…

20:00

we view the worst possible mistake as showing a bad note because that's going to undermine trust and the trust is is is why people…

20:30
#community-notes#quality#trust#moderation
Explainer1:13:30

Radical Transparency: Anyone Can Download the Code and Reproduce Every Note

Trust required total openness. The ranking code and all rating data are published so anyone can replicate the entire service and audit it. This forced awkward architectural choices — the matrix-factorization model is built to be re-run from a downloaded TSV — and it really is runnable: on one machine with ~500GB of RAM in about a day. Vitalik Buterin is among those who've verified it independently.

  • The code that decides which notes show, plus all ratings data, is public
  • Anyone can take the code and data and replicate the whole service
  • Being genuinely runnable forced unusual architectural decisions from scratch
  • Replicating it takes ~500GB RAM on one machine for about a day; Vitalik Buterin verified it in a blog post

the code that decides what note show has to be out in the open all the data and ratings that that make it happen have…

Keith Coleman · 1:14:00

oh like 500 gigs it it'll take like a day if you don't uh do anything special to speed it up

Jay Baxter · 1:16:30
#open-source#transparency#trust#audit

Story· 5

Story26:30

A Note Cuts Reposts 50-60% — and Makes Authors 80% More Likely to Delete

In an A/B test, showing a post with a note drove 30-40% drops in likes and reposts, versus the 1% effect size that's considered great for an algorithm change. External research groups found total reposts fall 50-60% once a note is applied, killing virality within a few generations. Authors also become 80% more likely to delete a noted post.

  • 1% is a strong effect size; notes produced 30-40% drops in likes and reposts in an A/B test
  • Independent research groups found ~50-60% drops in total reposts after a note is applied
  • At 50-60% per generation, virality quickly collapses to zero
  • Authors are 80% more likely to delete their post after being noted

1% is typically an awesome effect size for any sort of algorithm change we saw more like 30 to 40% engagement rate drops uh for…

27:00

authors become 80% more likely to decrease or sorry to delete their post after they get noted

28:30
#community-notes#virality#engagement#research
Story25:00

A Contributor With 12 Followers Made the White House Delete a Tweet

When Community Notes launched US-wide in 2022, a note appeared on a White House tweet and the White House deleted it and reissued an updated statement. Keith uses it to show the leverage an ordinary contributor has: someone with a handful of followers changed a government's public talking points.

  • A note on a 2022 White House tweet led the White House to delete and re-issue the statement
  • The note's author was an ordinary contributor, not a public figure
  • Demonstrates that anyone correct and credible can shape public discourse
  • Intrinsic impact, not pay, is what motivates top contributors

when we first launched us-wide this was like in 2022 a note appeared on a White House tweet and the White House deleted the tweet…

25:00

here you just put a put a note on the White House and they changed their public talking points based on what you did like…

25:30
#community-notes#impact#story#contributors
Story30:00

"How About I Just Stop Doing My Job" — The Origin of Community Notes

Keith joined Twitter in 2016 during the Trump-Clinton election, struck by how important and unsolved the misinformation problem was. After years running a large PM team without seeing the change he wanted, he came back from paternity leave and asked his boss Kavon to let him drop his job entirely to try crazy ideas on misleading information.

  • Keith joined Twitter in 2016 via acquisition during the presidential election
  • Existing approaches (fact-checkers, internal trust & safety) weren't working or trusted
  • He was weighing starting a company but kept returning to the misinformation problem
  • He asked his boss to let him quit his PM role to prototype; it became Community Notes

I went to my boss Kavon I was like Hey Kayon how about I just stop doing my job and I go work on this…

Keith Coleman · 33:00
#origin-story#career#product#misinformation
Story35:00

How Elon's DM Changed the Name Back From Bird Watch to Community Notes

The project was originally called Community Notes in Keith's very first Figma mockup, then renamed Bird Watch. After acquiring Twitter, Elon DM'd Keith calling it "this community notes thing" — the exact original name — and they changed it back the next day. Jack Dorsey mocked the name as the most boring Facebook-style name possible.

  • The first Figma mockup was already called Community Notes before it became Bird Watch
  • Elon kept referring to it as "community notes" in a post-acquisition DM
  • The name was switched back the next day
  • Descriptive, intuitive names have served the product well despite Jack's ribbing

he kept referring to it as this community notes thing and I was like you know it's interesting that you keep calling that calling it…

Keith Coleman · 35:30

Elon was like hey let's just call it that and so the next day we just changed the name

Keith Coleman · 36:00
#naming#product#elon-musk#story
Story1:18:30

The Israel-Hamas Deluge: The Biggest Stress Test — and Why Notes Held

The October 2023 Israel-Hamas conflict was the largest flood of misleading information Jay had ever seen, including out-of-context photos and fake battle footage made in the Arma 3 game engine. Features shipped just weeks earlier — media matching and a speed-up shaving hours off note time — proved decisive. Median time from post to note was about five hours, versus the two-to-four days typical of fact-checking.

  • Oct 2023 conflict produced an overwhelming volume of misleading posts and media
  • Fake battle footage was created in the Arma 3 video game simulator
  • Recently shipped image/video note-matching and a ~3-hour speed-up landed just in time
  • Median post-to-note time was ~5 hours vs. the 2-4 days typical of traditional fact-checking

there were people making fake like battle footage in the video game simulator Arma 3 so there like notes explaining the stuff looked really looked…

1:19:30

the median time from a post going live to a note showing up was five hours which is like crazy fast typical factchecking is like…

Keith Coleman · 1:21:00
#community-notes#misinformation#stress-test#speed

Takeaway· 2

Takeaway13:30

The Scale: 95,000 Notes, 30 Billion Views, Nearly a Million Contributors

Keith shares the numbers that surprise people. In 2024 roughly 95,000 notes were seen about 30 billion times, more than double the prior year. A single note on an image or video is auto-matched to every post containing that media, so one note can cover thousands of posts. Around 950,000 contributors now participate worldwide.

  • 2024: ~95,000 notes seen ~30 billion times, more than double 2023's ~37k notes / 14 billion views
  • By contrast a UC Berkeley figure cited ~10 traditional fact checks a day versus hundreds of notes a day
  • One note is matched to all posts sharing the same image or video, covering thousands of posts
  • ~950,000 contributors globally, with more on a waitlist

we had something like 95,000 notes that were seeing about 30 billion times that's more than double the prior year prior year was something like…

14:30

there's something like 950,000 contributors around the world that's you know nearing a million people making this happen

15:00
#community-notes#scale#metrics#fact-checking
Takeaway56:30

Let People Opt In: Self-Selection Beats Assigning a Team

No one was assigned to Community Notes — every member reached out or applied and was interviewed for mutual fit, so everyone was fully bought into the mission. Keith notes Elon later applied the same principle at company scale with the "fork in the road" email that asked all of Twitter to click a button to opt into the hardcore Twitter 2.0.

  • No one was assigned to the project; contributors self-selected and were interviewed
  • Self-selection produces total buy-in to the goal, team, and way of working
  • Elon's "fork in the road" email asked employees to actively opt into Twitter 2.0
  • Opt-in works at the start of a crazy project and, surprisingly, at large scale too

people are self- selecting to join we did not assign anyone to this project like people reached out to join or they applied to join

Keith Coleman · 56:30

he sent an email out that was like Hey Twitter 2.0 like Fork fork in the road right fork in the road fork in the…

Keith Coleman · 57:30
#team-building#hiring#culture#motivation