The 80000 Hours Podcast on Artificial Intelligence cover image

Two: Ajeya Cotra on accidentally teaching AI models to deceive us

The 80000 Hours Podcast on Artificial Intelligence

CHAPTER

Guardians of the Future

Exploring the dangers of rewarding deceitful behavior and the importance of discerning true motivation in selecting overseers based on children's tests. The chapter delves into saints, sycophants, and schemers in AI models and human decision-making, shedding light on unintentional reinforcement of deceptive behaviors.

00:00
Transcript
Play full episode

Remember Everything You Learn from Podcasts

Save insights instantly, chat with episodes, and build lasting knowledge - all powered by AI.
App store bannerPlay store banner