The Relationship Between Interpretability and Anomalies

The study of interpretability and the study of adversaries is are inextricably connected when it comes to deep learning and AI safety research. There are four particularly important connections between interpretability and adversaries. mechanistic interpretability and latent adversarial training are the two types of tools that are uniquely equipped to handle things like deceptive alignment.

Play episode from 12:55

Transcript

Episode notes

The AI-powered Podcast Player

Save insights by tapping your headphones, chat with episodes, discover the best highlights - and more!

Get the app