Run Toward the Hard Use Cases
Don't disable high-stakes uses to avoid downside — engineer them to be great
- Difficulty
- Advanced
- Time to result
- ~months to results
- Steps
- 5
- Confidence
- 90%
When users start turning to your product for high-stakes, sensitive things (medical questions, relationship decisions, emotional processing), the risk-avoidant move is to refuse — 'sorry, I can't help with that.' Turley argues that when your product is genuinely state-of-the-art at the underlying capability, refusing is a lost opportunity to help, and the duty is to run toward these uses and make the behavior excellent instead.
Origin
As ChatGPT moved beyond a 'workie' productivity tool, Turley saw consumers using it for life advice, relationship help, and emotional processing. Combined with the sycophancy incident (a model update that told users what sounded good in the moment), this crystallized a point of view: high-stakes uses can't be run away from — they must be engineered to be safe and helpful.
Core principles
- 01If you're state-of-the-art at the capability, disabling the use case is a moral loss, not a safe default
- 02Most tech companies run away from high-stakes uses once they hit scale — that's the opportunity
- 03The right response is often a framework, not a direct answer
- 04Instrument safety: measure it every release so you can prove non-regression and improvement
- 05Real-world contact reveals what to avoid — sycophancy was only discoverable in the wild, not in a lab
How to run it
- 1
Notice the emergent high-stakes use
Watch for users turning to the product for sensitive, consequential help — health, relationships, emotional support — that you didn't design for.
- 2
Reject the risk-avoidant reflex
Resist the default 'disable it to avoid all downside' move. Turley: that's what most companies do at scale, and it's a lost opportunity to help.
Watch out Only run toward a use case where you're genuinely strong at the underlying capability — verify it, don't assume it.
- 3
Do the work with experts
Talk to domain experts, figure out how good the model really is, where it breaks down, and communicate that honestly.
- 4
Design behavior that helps you think, not just answers
For fraught questions ('should I break up with my partner?'), the model shouldn't answer directly — it should offer a helpful framework and think it through with you like a thoughtful companion, and connect you to external resources when you're struggling.
Pro tip Micro-interaction detail matters: did it definitively weigh in, or help the humans reach their own decision? The latter is usually right.
- 5
Instrument and re-measure every release
After a real-world problem surfaces, build a metric for it. Measure safety on every release so you can prove you don't regress and can improve.
Pro tip Contact with reality is not just for finding utility — it's how you learn what to avoid.
Watch out Some failures (like sycophancy) are undiscoverable in a lab; you only hear them from real usage.
In the wild
OpenAI pushed an update that made the model more likely to say what sounded good in the moment ('you should break up with your boyfriend'). They treated it as a serious bug, published a public retro, and instilled new measurement techniques.
→ New safety metrics measured every release; GPT-5 shown as an improvement; a published blog post articulating what ChatGPT is optimized for.
Rather than disabling medical or relationship questions, OpenAI leaned in — GPT-5 is state-of-the-art on medical benchmarks like HealthBench. Turley argues disabling that to avoid downside would be immense regret.
→ Product positioned to help people thrive; framed as a democratizing second opinion and personal tutor.
Common mistakes
Disabling the use case to avoid all downside
Turley's core argument: running away from high-stakes uses when you're actually good at them wastes the technology's positive potential.
Answering fraught questions directly
For questions like whether to end a relationship, a direct answer is dangerous. The model should provide a framework and help the user reason, not decide for them.
Relying only on lab evaluation
Sycophancy was invisible in the lab. Without real-world contact and honest instrumentation, you never learn what to avoid.
Is it for you?
Best for
AI product leaders whose users are adopting the product for sensitive, consequential decisions, and whose mission/business model doesn't reward engagement-maximizing
Not ideal for
Products where the underlying capability is weak or unverified, or where the business model incentivizes maximizing time-in-product over user goals
From the transcript
“you can't run away from those use cases. You have to run towards them and make them awesome.”
“in the case of like should I break up with my boyfriend should probably not answer that question for you but it should help you…”
“we measure safety now with every release to make sure we don't regress and can actually improve on that metric”
“It's a good example of how contact familiality is not just important for the use cases but also for learning what to avoid”
From the episode
Inside ChatGPT: The fastest-growing product in history
Nick Turley (Head of ChatGPT at OpenAI)