"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis cover image

Dodging Latent Space Detectors: Obfuscated Activation Attacks with Luke, Erik, and Scott.

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

00:00

Backdoor Attacks and Defense Mechanisms in AI

This chapter investigates the security vulnerabilities of machine learning models, particularly focusing on backdoor attacks initiated through malicious data manipulation. It emphasizes methodologies for detecting harmful behaviors, including the use of logistic regression and various detection strategies to distinguish between benign and harmful requests. The discussion also highlights the complexities of model access during attacks, detailing different access scenarios and the challenges in identifying and mitigating backdoor threats.

Transcript
Play full episode

The AI-powered Podcast Player

Save insights by tapping your headphones, chat with episodes, discover the best highlights - and more!
App store bannerPlay store banner
Get the app