Interconnects cover image

Interviewing Louis Castricato of Synth Labs and Eleuther AI on RLHF, Gemini Drama, DPO, founding Carper AI, preference data, reward models, and everything in between

Interconnects

00:00

Challenges of Algorithm Distillation in NLP and Reinforcement Learning

This chapter explores the intricacies of model-based reinforcement learning, focusing on the computational challenges of muesli and its superiority over AlphaGo. It also addresses the difficulties of applying algorithm distillation to NLP, revealing key limitations of current models and their attention mechanisms in complex tasks.

Play episode from 19:19
Transcript

The AI-powered Podcast Player

Save insights by tapping your headphones, chat with episodes, discover the best highlights - and more!
App store bannerPlay store banner
Get the app