
33 - RLHF Problems with Scott Emmons
AXRP - the AI X-risk Research Podcast
00:00
Intro
Exploring challenges of AI deception and partial observability in reinforcement learning from human feedback, discussing theoretical concerns and real-world implications for deployed AI systems like chat GPT and llama.
Transcript
Play full episode