“how fast we're even saturating Sweetbench once we focus on it.”
SWE-bench
0 recommend/use · 2 sourced episodes
Every sourced reference
Short attributed excerpts only. Timestamps are approximate.
“we're at 50% on SUB bench which is this like you know benchmark around how well the models are at coding.”
Related frameworks
Build Only at the Magic Intersection
Don't ship what anyone could build off the shelf — build only where model and product uniquely meet.
Build-With-The-Tools Assessment
Don't test people on doing work without AI — hand them the tools and score what they can build in an hour
Defensible Moats for AI Startups
Four durable places to build in AI where foundation-model labs are least likely to squash you.
Elastic-Demand Career Bet
Invest in domains where making people 10x more productive increases demand rather than reducing it
Hunting New Bottlenecks When AI Writes the Code
When AI removes the coding bottleneck, constraints shift up and downstream — go find them.
Make the Other Mistake (Prompting for Brutal Feedback)
To get honest AI critique, over-correct toward brutal — the model won't actually overshoot.
Patient-Then-Floor-It Hiring
Be obsessively patient hiring the founding team; step on the gas the moment demand exceeds what you can handle
Pull-Based Product-Market Fit
Be stubborn on your thesis but open on its form — chase the customer who is surprisingly easy to sell, not the one you must force
RLAIF Reward Design
Have an expert define success criteria and a rubric once, then let AI reinforce the capability — more scalable than labeling examples
The 10-to-1 Input/Output Kill Test
When you pour 10 units of effort in for 1 unit of output, the project has run its course.
The Three-Part AI Product Utility Equation
A useful AI product needs model intelligence, context/memory, and application/UI to all converge.
Three Tenets: Can-Do, High Standards, Intensity
The three cultural tenets Foody credits for the fastest revenue ascent in history
Value-Chain Eval
Before applying AI to your business, build a systematic test that measures how well it automates your core value chain
People in these episodes
Related resources
- AI 2027
- AlphaSights
- Anthropic
- Anthropic Console Workbench
- Apple Podcasts
- Benchmark
- Bolt
- ChatGPT
- Claude
- Claude Code
- Claude Opus 4
- Codex
- Cursor
- Enterpret
- General Catalyst
- GLG
- Goldman Sachs
- Google Gemini
- Google Reader
- GPQA
- Handshake
- Harvey
- High Output Management
- Jira Product Discovery
- Lenny's Newsletter
- Lenny's Podcast
- Lovable
- McKinsey & Company
- Microsoft Windows
- Model Context Protocol
- Neuralink
- No Priors
- OneSchema FileFeeds 2.0
- OpenAI
- Oppenheimer
- Productboard
- Prompt Improver
- Scale AI
- Shoe Dog
- Shopify
- Spotify
- Stripe
- Suits
- Surge AI
- Tesla
- The Goal
- Uber
- Warrant
- WorkOS
- xAI
- Zanzibar
- Zero to One
Spot an error or want this page removed? Request a correction or removal.