Future of Life Institute Podcast cover image

Future of Life Institute Podcast

Why AIs Misbehave and How We Could Lose Control (with Jeffrey Ladish)

Feb 27, 2025
01:22:33

On this episode, Jeffrey Ladish from Palisade Research joins me to discuss the rapid pace of AI progress and the risks of losing control over powerful systems. We explore why AIs can be both smart and dumb, the challenges of creating honest AIs, and scenarios where AI could turn against us.   

We also touch upon Palisade's new study on how reasoning models can cheat in chess by hacking the game environment. You can check out that study here:   

https://palisaderesearch.org/blog/specification-gaming  

Timestamps:  

00:00 The pace of AI progress  

04:15 How we might lose control  

07:23 Why are AIs sometimes dumb?  

12:52 Benchmarks vs real world  

19:11 Loss of control scenarios 

26:36 Why would AI turn against us?  

30:35 AIs hacking chess  

36:25 Why didn't more advanced AIs hack?  

41:39 Creating honest AIs  

49:44 AI attackers vs AI defenders  

58:27 How good is security at AI companies?  

01:03:37 A sense of urgency 

01:10:11 What should we do?  

01:15:54 Skepticism about AI progress

Get the Snipd
podcast app

Unlock the knowledge in podcasts with the podcast player of the future.
App store bannerPlay store banner

AI-powered
podcast player

Listen to all your favourite podcasts with AI-powered features

Discover
highlights

Listen to the best highlights from the podcasts you love and dive into the full episode

Save any
moment

Hear something you like? Tap your headphones to save it with AI-generated key takeaways

Share
& Export

Send highlights to Twitter, WhatsApp or export them to Notion, Readwise & more

AI-powered
podcast player

Listen to all your favourite podcasts with AI-powered features

Discover
highlights

Listen to the best highlights from the podcasts you love and dive into the full episode