AXRP - the AI X-risk Research Podcast

20 - 'Reform' AI Alignment with Scott Aaronson

18 snips
Apr 12, 2023
Ask episode
AI Snips
Chapters
Transcript
Episode notes
ADVICE

Use Concentration Of Power To Deploy Safety Standards

  • When few players control state-of-the-art models, push for industry-wide technical mitigations (e.g., watermarking) to become standards.
  • Aaronson notes high training costs give a window where convincing leading labs can achieve broad effect.
INSIGHT

Focus On Detecting Model Origin As A Tractable Defense

  • Aaronson anticipated mass misuse patterns (impersonation, essay cheating, propaganda) and focused on detecting model-origin as a tractable defensive target.
  • He pivoted from abstract alignment to watermarking because it yields provable, implementable defenses for current harms.
INSIGHT

Undetectable Watermarking Without Quality Loss

  • A statistical watermark can be inserted by pseudorandomly biasing token sampling without degrading output quality.
  • Aaronson gives a simple rule using pseudorandom scores r_i and model probabilities p_i so sampling remains statistically identical yet detectable later.
Get the Snipd Podcast app to discover more snips from this episode
Get the app