Chris Olah
Co-founder and Interpretability Research Lead, Anthropic
Leads Anthropic's interpretability research, the effort to reverse-engineer what happens inside neural networks.
Chris Olah co-founded Anthropic and leads its interpretability research, which tries to work out, neuron by neuron and circuit by circuit, what a trained model is actually computing. His own site describes the work as reverse engineering neural networks "into human understandable algorithms."
He has no university degree. Raised in Toronto, he left the University of Toronto after about a year, and in 2012 received a 100,000 dollar Thiel Fellowship, which pays people under 20 to skip or leave college. He told the 80,000 Hours podcast that the adults around him "totally came around" once someone had given him the money. His blog posts explaining neural networks with diagrams became standard teaching material.
He started at Google Brain as an intern in 2015 and stayed three years as a research scientist, co-writing the DeepDream work and feature-visualization research that showed what individual neurons respond to. In 2017 he co-founded Distill, an online journal for clearly explained machine-learning research, and edited it until it went on hiatus in 2021. In 2018 he moved to OpenAI to lead interpretability, where his team published the Circuits project and found "multimodal neurons" in the CLIP model that fired for a concept such as Spider-Man whether it appeared as a photo, a drawing or a word (published in March 2021). He left with Dario Amodei's group at the end of 2020.
At Anthropic his team applied dictionary learning to language models, extracting millions of interpretable features from Claude 3 Sonnet in May 2024; amplifying one of them produced "Golden Gate Claude," a model that steered every conversation toward the bridge. In May 2026 he spoke at the Vatican presentation of Pope Leo XIV's encyclical on AI, arguing that outsiders must oversee companies like his own.
Known for
Circuits
The 2020 Distill series, opened by "Zoom In" (March 2020), argued that neural networks contain meaningful features connected into readable circuits that can be studied like biology, setting the program for mechanistic interpretability.
Scaling Monosemanticity
The May 2024 Anthropic paper used sparse dictionary learning to pull millions of features out of a production model, Claude 3 Sonnet, including ones for deception and for the Golden Gate Bridge, and showed that turning them up or down changed the model's behavior.
Distill
He co-founded the journal in 2017 to publish interactive, carefully explained research; it paused new publications in 2021.
On the record
He argues that AI companies, including his own, are shaped by commercial and geopolitical incentives and so need oversight from people outside the industry.
“No matter how sincerely any of us intend to do the right thing, and I believe many of us do, we will always be influenced by those incentives”
Prepared remarks at the Vatican, quoted by Fortune, June 2026, 2026
Career
Sources
- Leadership at Anthropic (Anthropic, checked September 2026)
- Christopher Olah, About (colah.github.io)
- Chris Olah on working at top AI labs without an undergrad degree (80,000 Hours podcast #108, 2021)
- Who is Chris Olah? (Fortune, June 3, 2026)
- Zoom In: An Introduction to Circuits (Distill, March 2020)
- Multimodal Neurons in Artificial Neural Networks (Distill, March 2021)
- Scaling Monosemanticity (Anthropic, Transformer Circuits Thread, May 2024)
- Inceptionism: Going Deeper into Neural Networks (Google Research, June 2015)
Brief profiles cover a person's professional record: roles, work and public statements, each from a source listed here. Full profiles, with positions across the debates, are kept for the people the site follows most closely.