
[HUMAN VOICE] "How useful is mechanistic interpretability?" by ryan_greenblatt, Neel Nanda, Buck, habryka
LessWrong (Curated & Popular)
00:00
Exploring Model Safety and Ablations
This chapter discusses the speakers' doubts about the effectiveness of enumerative safety and the challenges posed by superhuman models. It also explores the concept of ablations and proposes retraining the model and augmenting smaller models with explanations for better performance.
Transcript
Play full episode