
Kopec Explains Software #128 Copyright & Machine Learning Models
Dec 11, 2023
Discussion of whether training machine learning models on copyrighted text and images is legally permissible. Exploration of expressive versus functional use and the four-factor fair use test. Debate over data‑mining safe harbor proposals and how legal uncertainty advantages large firms. Consideration of international rules, artist concerns, and when AI outputs qualify for copyright.
AI Snips
Chapters
Transcript
Episode notes
Expressive Purpose Determines Infringement
- Copyright infringement requires using a work for its expressive purpose, not merely copying its material form.
- Kopec cites Baker v. Selden to argue ML training can be non-expressive if models only extract ideas, not the original expression.
Fair Use Is The Core Legal Test For ML Training
- U.S. courts evaluate reuse of copyrighted material under four fair use factors when deciding ML training legality.
- Kopec lists purpose/character, nature, amount/substantiality, and market effect as the weighing criteria in cases.
Fair Use Litigation Favors Big Tech
- Current fair use litigation favors organizations that can afford legal defense, disadvantaging small innovators.
- Kopec notes big companies like OpenAI or Google can bankroll court tests, creating asymmetric risk for startups.
