Mapping the Mind of a Neural Net: Goodfire’s Eric Ho on the Future of Interpretability

150 snips

Jul 8, 2025

Eric Ho, founder of Goodfire, is at the forefront of AI interpretability, tackling the challenge of understanding neural networks. He shares breakthroughs in resolving superposition using sparse autoencoders and demonstrates innovative model editing techniques. The conversation touches on real-world applications, particularly in genomics, and the vital role of interpretability as AI grows in influence. Ho emphasizes the importance of independent research in making AI systems more transparent and mitigating risks associated with powerful technologies.

Ask episode

AI Snips

Chapters

Transcript

Episode notes

INSIGHT

Necessity of Neural Net Interpretability

Understanding neural networks is crucial as AI gains mission-critical societal roles.
Inspecting the model's inner workings enhances safety, power, and reliable AI design beyond black-box evaluation.

INSIGHT

Advantage in Mechanistic Interpretability

Mechanistic interpretability benefits from perfect access to all neural network parameters, unlike neuroscience.
This advantage allows deeper progress in mapping and understanding large language models' inner workings.

INSIGHT

Bonsai Metaphor for AI Design

Goodfire aims to shape AI like bonsai trees by intentionally growing and pruning behaviors.
Understanding every piece of training data's effect enables precise, deliberate AI cognition design.

Get the Snipd Podcast app to discover more snips from this episode

Get the app

Eric Ho is building Goodfire to solve one of AI’s most critical challenges: understanding what’s actually happening inside neural networks. His team is developing techniques to understand, audit and edit neural networks at the feature level. Eric discusses breakthrough results in resolving superposition through sparse autoencoders, successful model editing demonstrations and real-world applications in genomics with Arc Institute's DNA foundation models. He argues that interpretability will be critical as AI systems become more powerful and take on mission-critical roles in society.

Hosted by Sonya Huang and Roelof Botha, Sequoia Capital

Mentioned in this episode:

Mech interp: Mechanistic interpretability, list of important papers here
Phineas Gage: 19th century railway engineer who lost most of his brain’s left frontal lobe in an accident. Became a famous case study in neuroscience.
Human Genome Project: Effort from 1990-2003 to generate the first sequence of the human genome which accelerated the study of human biology
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
Zoom In: An Introduction to Circuits: First important mechanistic interpretability paper from OpenAI in 2020
Superposition: Concept from physics applied to interpretability that allows neural networks to simulate larger networks (e.g. more concepts than neurons)
Apollo Research: AI safety company that designs AI model evaluations and conducts interpretability research
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. 2023 Anthropic paper that uses a sparse autoencoder to extract interpretable features; followed by Scaling Monosemanticity
Under the Hood of a Reasoning Model: 2025 Goodfire paper that interprets DeepSeek’s reasoning model R1
Auto-interpretability: The ability to use LLMs to automatically write explanations for the behavior of neurons in LLMs
Interpreting Evo 2: Arc Institute's Next-Generation Genomic Foundation Model. (see episode with Arc co-founder Patrick Hsu)
Paint with Ember: Canvas interface from Goodfire that lets you steer an LLM’s visual output in real time (paper here)
Model diffing: Interpreting how a model differs from checkpoint to checkpoint during finetuning
Feature steering: The ability to change the style of LLM output by up or down weighting features (e.g. talking like a pirate vs factual information about the Andromeda Galaxy)
Weight based interpretability: Method for directly decomposing neural network parameters into mechanistic components, instead of using features
The Urgency of Interpretability: Essay by Anthropic founder Dario Amodei

On the Biology of a Large Language Model: Goodfire collaboration with Anthropic