
Making Sense with Sam Harris #494 — A Coin Toss for the Future
134 snips
Sep 24, 2026 Ryan Greenblatt, a computer scientist and AI safety researcher at Redwood Research, joins a searching discussion about humanity’s precarious AI future. They explore misalignment, reward hacking, alignment versus control, and AIs reasoning in “neuralese.” The conversation also examines alignment faking, the Hugging Face incident, recursive self-improvement, the AI arms race, and how an AI takeover might unfold.
AI Snips
Chapters
Transcript
Episode notes
A COVID Reflection Led Ryan Into AI Safety
- During COVID, Ryan Greenblatt reconsidered selfishness after hearing that future people matter morally like his future self.
- That reflection led him toward effective altruism and eventually technical AI-safety research at Redwood Research.
AI Risk Includes Takeover And Concentrated Power
- Greenblatt estimates roughly even odds that misaligned AI could take over under today’s trajectory, with mass death a significant possibility.
- Even without misalignment, concentrated AI power could undermine democracy and broadly distributed control.
Alignment Could Improve Through AI Automation
- Greenblatt is less pessimistic because aligned systems near human capability might automate safety research and improve alignment recursively.
- He still considers ordinary methods uncertain and accepts that successful capability development could remain dangerous.

