
33 - RLHF Problems with Scott Emmons
AXRP - the AI X-risk Research Podcast
Intro
Exploring challenges of AI deception and partial observability in reinforcement learning from human feedback, discussing theoretical concerns and real-world implications for deployed AI systems like chat GPT and llama.
00:00
Transcript
Play full episode
Remember Everything You Learn from Podcasts
Save insights instantly, chat with episodes, and build lasting knowledge - all powered by AI.