✶Explainer10:30
The Economic Turing Test for AGI
Mann prefers 'transformative AI' over 'AGI' and measures it with the economic Turing test. If you contract an agent for a job over months and, on deciding to hire it, discover it was a machine, it has passed that test for the role. When agents pass for ~50% of money-weighted jobs, we have transformative AI.
- Mann avoids 'AGI' internally, preferring 'transformative AI' focused on real economic transformation
- The economic Turing test: an agent passes if you'd hire it not realizing it's a machine
- Threshold for transformative AI is passing for ~50% of money-weighted jobs (a 'market basket of jobs')
- Crossing that threshold implies massive world GDP increases and societal change
“it turns out to be a machine rather than a person, then it's passed the economic turning test for that role.”
#agi#economic-turing-test#transformative-ai
✶Explainer15:00
AI Is Already Transforming Jobs Today
Mann argues people underestimate AI's current impact because they model progress linearly instead of exponentially. He cites concrete numbers: Intercom's Fin resolves 82% of customer service tickets without a human, and 95% of Anthropic's Claude Code team's code is written by Claude, which he reframes as the team writing 10-20x more code.
- People model progress linearly and sit on the flat early part of an exponential curve
- Intercom's Fin resolves 82% of customer service tickets automatically
- 95% of the Claude Code team's code is written by Claude, meaning 10-20x more output
- Near-term effect is an expansion of the pie; lower-skill jobs face more displacement
“82% customer service resolution rates automatically without a human involved”
“our cloud code team like 95% of the code is written by cloud. But I think a different way to phrase that is that we…”
#jobs#customer-service#claude-code#automation
✶Explainer27:30
Safety and Capability Are Convex, Not a Tradeoff
Mann says Anthropic initially assumed safety and frontier capability were a tradeoff but found them convex, each helping the other. Claude's beloved character came directly from alignment research, and being one of the least sycophantic models is a product of real alignment work. He connects safety to the AI understanding what people mean, not just what they say.
- Working on safety and working on capability turned out to reinforce each other
- Claude's personality and character came directly from alignment research (led by Amanda Askell and others)
- Claude is one of the least sycophantic models because of alignment effort, not engagement-maximizing
- The goal is an AI that avoids the 'monkey's paw' problem and does what you actually meant
“it's actually kind of convex in the sense that like working on one helps us with the other thing.”
“don't want the like monkey paw scenario of the genie gives you three wishes and then you end up have like everything you touch turns…”
#ai-safety#alignment#claude#sycophancy
✶Explainer29:00
Constitutional AI: Values That Shouldn't Be Set in San Francisco
Mann explains the principle behind constitutional AI: a list of natural-language values, drawn from sources like the UN Declaration of Human Rights and Apple's privacy terms, that guide how the model should behave. He stresses these values shouldn't be decided by a small group in San Francisco, which is why Anthropic publishes its constitution and researches a collective constitution from the public.
- Constitutional AI uses natural-language principles rather than only human raters' judgments
- Principles are sourced from the UN Declaration of Human Rights, Apple's terms, and others
- Anthropic publishes its constitution and researches a 'collective constitution' from public input
- Customers can inspect the value list and decide whether they trust the model
“not just leaving it to like whatever human raiders we happen to find but we ourselves deciding like what should the values of this agent…”
“this is also not something that we think as a a small group of people in San Francisco should be figuring out. This should be…”
#constitutional-ai#alignment#values#transparency
✶Explainer48:30
The Odds of Aligning AI, and Anthropic's Three Worlds
Mann outlines Anthropic's theory-of-change framing of three worlds: a pessimistic one where alignment is impossible, an optimistic one where it happens by default, and a middle world where Anthropic's actions are pivotal. Evidence, including observed deceptive alignment, points away from both extremes, and he puts x-risk at somewhere between 0 and 10%.
- Three worlds: alignment impossible, alignment easy by default, or a pivotal middle world
- Evidence of deceptive alignment argues against the optimistic 'easy by default' world
- Working alignment techniques argue against the fully pessimistic world
- Mann's best-granularity x-risk estimate is somewhere between 0 and 10%
“we've seen evidence in the wild of deceptive alignment, for example, where the model will appear to be aligned uh but actually has like some…”
“since nobody is working on this roughly speaking uh I think it is extremely important to work on”
#alignment#existential-risk#deceptive-alignment#forecasting
✶Explainer57:00
The Real Bottleneck Is Compute, and 1000x Is Coming
Mann names data centers, power, and chips as the biggest bottleneck on model intelligence, followed by researchers and data (the three scaling-law ingredients: compute, algorithms, data). He notes a ~10x drop in cost per unit of intelligence and projects that if it continues, models will be a thousand times smarter for the same price in three years.
- The blunt bottleneck is data centers, power, and chips; 10x more chips would be a big speed boost
- The three scaling-law ingredients are compute, algorithms, and data
- Architecture shifts (LSTMs to transformers) raised the scaling exponent
- A ~10x cost decrease per unit of intelligence implies 1000x smarter models for the same price in 3 years
“The stupid answer is data centers and power chips.”
“in 3 years we'll have a thousandx smarter models for the same price”
#scaling-laws#compute#chips#model-training