261 episodes
- Tom McGrath is co-founder and Chief Scientist at Goodfire, and a former Google DeepMind researcher. He joins Tim Scarfe to ask what neural networks actually learn, whether their internal representations converge on structures in the world, and whether interpretability can extract new scientific knowledge rather than merely explain model outputs.
Beginning with AlphaZero and learned modularity, the conversation moves into neural geometry: concept manifolds, reusable computation inside Llama, and why activation steering can fail when it pushes a model off-manifold. McGrath then makes the case for intentional design, using interpretability as part of the training loop. They examine controlled generalisation, features as rewards, predictive data debugging, and the uncomfortable fact that a model may recognise a hallucination or reward hack and still produce it.
The discussion closes on grader awareness, oversight and collusion between adaptive agents, then returns to sparse autoencoders. SAEs are useful, McGrath argues, but they may fracture the higher-dimensional structures networks actually use. This episode was made with support from Goodfire.
---
TIMESTAMPS:
00:00:00 Introduction: Can interpretability speed-run science?
00:02:03 The invisible grader
00:06:51 What AlphaZero learned from the world
00:12:24 Interpretability as a control loop
00:21:54 The forbidden method and safer interventions
00:37:36 Why models catch hallucinations too late
00:46:19 Debug the dataset before training
00:50:44 Why neural networks become modular
00:55:57 Finding the geometry inside a network
01:02:55 Why steering falls off the manifold
01:12:10 A reusable calculator inside Llama
01:17:19 From abstractions to goals
01:25:28 Reward hacking, oversight and collusion
01:37:23 Are sparse autoencoders dead?
---
REFERENCES:
paper:
[00:05:45] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
https://arxiv.org/abs/2502.17424v7
[00:11:05] Acquisition of Chess Knowledge in AlphaZero
https://arxiv.org/abs/2111.09259
[00:25:30] Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
https://arxiv.org/abs/2507.16795
[00:29:30] Persona Vectors: Monitoring and Controlling Character Traits in Language Models
https://arxiv.org/abs/2507.21509
[00:41:14] Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
https://arxiv.org/abs/2602.10067
[00:47:03] Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal
https://arxiv.org/abs/2606.12360
[01:00:26] Do Sparse Autoencoders Capture Concept Manifolds?
https://arxiv.org/abs/2604.28119
[01:03:04] Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
https://arxiv.org/abs/2605.05115
[01:14:20] Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
https://arxiv.org/abs/2605.01148
[01:29:35] Measuring Reward-Seeking via Contrastive Belief Updates
https://arxiv.org/abs/2607.18966v1
other:
[00:15:44] Intentional Design
https://www.goodfire.com/blog/intentional-design
[00:56:12] The World Inside Neural Networks
https://www.goodfire.com/research/the-world-inside-neural-networks
[01:37:28] A Pragmatic Vision for Interpretability
https://www.alignmentforum.org/posts/StENzDcD3kpfGJssR/a-pragmatic-vision-for-interpretability
---
RESCRIPT:
https://app.rescript.info/share/846cfee4131b664fd09209cc3b98018e Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander Panfilov
22/08/2026 | 49 mins.Tim Scarfe speaks with Ilia Shumailov and Alexander Panfilov about their paper, Stealing Reasoning Traces from Proprietary LLM APIs.The core bug sounds deceptively simple: providers return encrypted reasoning state so conversations can be resumed or forked. But those blobs can be replayed across users and sibling models. A smaller model can ask the provider to decrypt the trace, then repeat the hidden reasoning in plain text. The discussion covers leaked private data, a broadly reusable jailbreak, poisoned agent traces, chain-of-thought monitoring, responsible disclosure, and possible defenses.Ilia Shumailov is an AI and security researcher, formerly at Google DeepMind, who completed his Cambridge PhD under Ross Anderson. Alexander Panfilov is a PhD researcher at the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems, working on AI safety, adversarial machine learning, and LLM red-teaming. They close by separating the demonstrated jailbreaking threat from ordinary benign distillation, and by arguing for controlled experiments over sweeping claims.---TIMESTAMPS:00:00:00 Intro montage00:01:33 Portable encrypted thought and decoded reasoning00:24:55 How the attack works and what it means00:39:04 Doom, defense, and scientific restraint---REFERENCES:paper:[00:00:00] Stealing Reasoning Traces from Proprietary LLM APIshttps://arxiv.org/abs/2608.09867[00:09:22] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyhttps://arxiv.org/abs/2507.11473[00:11:30] Reasoning Models Don’t Always Say What They Thinkhttps://www.anthropic.com/research/reasoning-models-dont-say-think[00:37:22] PostTrainBench: Can LLM Agents Automate LLM Post-Training?https://arxiv.org/abs/2603.08640[00:41:02] Large-scale online deanonymization with LLMshttps://arxiv.org/abs/2602.16800other:[00:09:28] OpenAI and Hugging Face partner to address security incident during model evaluationhttps://openai.com/index/hugging-face-model-evaluation-security-incident/[00:10:22] Claude, GPT, and Gemini All Struggle to Evade Monitorshttps://metr.org/notes/2025-08-22-claude-gpt-gemini-struggle-evade-monitors/tool:[00:42:08] Isabelle proof assistanthttps://isabelle.in.tum.de/---RESCRIPT: https://app.rescript.info/share/07fc38276e0823dc9b8986c32e202c7f- Astrophysicist Adam Becker, author of "What Is Real?", joins Tim Scarfe to take apart the futures Silicon Valley keeps selling: the 2045 singularity, mind uploading, Mars colonies, and the AI apocalypse. His new book *More Everything Forever* argues these ideas are hugely influential, mostly evidence-free, and bankrolled by tech billionaires who need a story in which growth never ends.Becker does the physics the boosters skip. Kurzweil's "law of accelerating returns" rests on cherry-picked data, and every exponential ends. Grant Bezos his perpetual energy growth and humanity boils the oceans within a few centuries, then exhausts the observable universe in under 4,000 years. The stars are too far away, Mars dirt is poison, and the day the dinosaur-killing asteroid hit Earth was still nicer than any day on Mars. On AI, Becker calls LLMs pocket calculators for language: hallucination is the model doing exactly what it always does, and the intelligence explosion assumes intelligence is a single number you can buy with compute.The sting is that Becker thinks the doomers are sincere. Yudkowsky, Bostrom and the effective altruists are not grifters, he says, just wrong, and their warnings that AI could end the world feed the same growth story the money depends on. He closes with his own prescription: take social problems seriously, regulate the whole tech industry, and tax billionaires out of existence.---TIMESTAMPS:00:00:00 Cold open and the thesis of More Everything Forever00:04:24 Kurzweil's singularity and the physical limits of exponential growth00:14:02 High agency and the fantasy of imprinting humanity on the cosmos00:16:55 Mind uploading, functionalism, and embodied cognition00:24:24 AI psychosis and anthropomorphizing LLMs00:26:24 Calculators, hallucination, and the limits of scale00:32:20 Yudkowsky and the intelligence-explosion argument00:40:37 True believers, venture capital, and the sci-fi growth narrative00:47:21 From Extropians to EA: utilitarianism and longtermism00:53:50 Brain worms and Becker's prescription: take social science seriously00:56:49 Why the AI-ethics discourse is broken01:01:42 The eugenics and IQ argument against 'intelligence'01:06:07 Why space settlement fails: Mars, the moon, and orbital data centers01:10:42 Billionaire myths and the search for purpose01:13:38 Tax billionaires, regulate tech: closing prescriptions---REFERENCES:book:[00:00:07] More Everything Forever (Adam Becker, 2025)https://www.hachettebookgroup.com/titles/adam-becker/more-everything-forever/9781541619593/[00:00:15] What Is Real? (Adam Becker, 2018)https://en.wikipedia.org/wiki/What_Is_Real%3F[00:15:46] What We Owe the Future (Will MacAskill, 2022)https://www.hachettebookgroup.com/titles/william-macaskill/what-we-owe-the-future/9781541618626/other:[00:00:27] Dreaming Against the Machine (podcast)https://www.dreamingagainstthemachine.com[00:01:04] The Useful Idiots of AI Doomsaying (Adam Becker, The Atlantic, 2025)https://www.theatlantic.com/books/archive/2025/09/what-ais-doomers-and-utopians-have-in-common/684270/
RESCRIPT: https://app.rescript.info/share/d6e37f9866673d8f74a39076efa5926b - This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlstWhy can deep networks discover abstractions that shallow models miss? Statistical physicist Matthieu Wyart joins Tim Scarfe to argue that the answer lies in the hidden hierarchy of data. Language and images are built from parts within parts; depth lets a network recover those coarse-grained variables and escape the curse of dimensionality.The conversation moves from jamming transitions and rough loss surfaces to Chomsky, context-free grammars and machine creativity. Wyart explains why next-token prediction can still recover compositional structure, where current systems fall short of genuine scientific invention, and why predicting latent representations rather than raw tokens could make learning far more sample-efficient.They also examine diffusion models, neural scaling laws and the limits of physics-inspired theory. The final question is on a personal note: if mistakes are the price of leaving the beaten path, how much scientific risk is worth taking?---TIMESTAMPS:00:00:00 Can machines learn abstractions from data?00:02:00 Notion agentic workspace00:02:49 From statistical physics to machine learning00:06:40 What physics can explain about learning00:16:37 From Carnot to Chomsky bulldozer00:21:21 How deep networks recover hidden hierarchies00:32:43 Where machine creativity still falls short00:40:48 How deep nets escape the curse of dimensionality00:52:19 Why predict latents instead of tokens01:02:49 The sample-efficiency case for latent prediction01:08:31 Diffusion, scaling laws and text entropy01:16:40 The scientists we learn from and the mistakes we make---REFERENCES:person:[00:00:43] Noam Chomskyhttps://linguistics.mit.edu/user/chomsky/tool:[00:02:08] Notion Developer Platformhttps://www.notion.com/en-gb/blog/introducing-developer-platformpaper:[00:04:43] Mastering the game of Go with deep neural networks and tree searchhttps://www.nature.com/articles/nature16961[00:05:52] Reconciling modern machine-learning practice and the bias-variance trade-offhttps://arxiv.org/abs/1812.11118[00:25:54] How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Modelhttps://arxiv.org/abs/2307.02129[00:42:12] Efficient Estimation of Word Representations in Vector Spacehttps://arxiv.org/abs/1301.3781[00:52:46] Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecturehttps://arxiv.org/abs/2301.08243[00:52:54] Learn from your own latents and not from tokens: A sample-complexity theoryhttps://arxiv.org/abs/2605.27734[01:08:31] A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Datahttps://arxiv.org/abs/2402.16991[01:11:39] Scaling Laws for Neural Language Modelshttps://arxiv.org/abs/2001.08361[01:12:17] Deriving Neural Scaling Laws from the statistics of natural languagehttps://arxiv.org/abs/2602.07488[01:13:34] Prediction and Entropy of Printed Englishhttps://ieeexplore.ieee.org/document/6773263---LINKS:Download PDF transcript: https://app.rescript.info/share/f7644cdaa86c5cc1e41e484e290f2bd4
- Can an AI do the right thing for the wrong reason? Tim Scarfe speaks with Apollo Research’s Alexander Meinke, Axel Højmark and Jérémy Scheurer about Measuring Reward-Seeking via Contrastive Belief Updates, their new research with OpenAI.
The panel asks how models infer what graders reward, why good behaviour can come from the wrong reason, and whether that difference can be measured. The conversation moves through promise-breaking, grader awareness, reward hacking, scheming, opaque reasoning and corrigibility, then turns to a detailed walkthrough of the contrastive-belief method and what its results do and do not show. The o3 results discussed here concern an intermediate checkpoint without safety training.
This episode was made in partnership with Apollo Research. MLST retained full editorial control.
Reference
Apollo Research: https://www.apolloresearch.ai/
---
TIMESTAMPS:
00:00:00 Cold Open
00:02:12 Right Things, Wrong Reasons
00:12:47 Grader Awareness
00:26:22 Legibility
00:32:35 What To Call It
00:35:58 Intelligence, Agency, Anthropomorphism
00:45:16 Apollo’s Mission
00:48:54 The End of the Exponential
00:55:45 The Paper
01:16:34 Closing Reflection
---
REFERENCES:
tool:
[00:00:08] Claude Fable
https://www.anthropic.com/claude/fable
[00:12:50] AlphaGo Zero
https://deepmind.google/blog/alphago-zero-starting-from-scratch/
[00:44:30] AlphaFold 3
https://deepmind.google/science/alphafold/
paper:
[00:01:02] Measuring Reward-Seeking via Contrastive Belief Updates
https://arxiv.org/abs/2607.18966
[00:16:19] Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
https://transformer-circuits.pub/2026/nla/
[00:26:48] Stress Testing Deliberative Alignment for Anti-Scheming Training
https://arxiv.org/abs/2509.15541
[00:35:33] Shortcut learning in deep neural networks
https://arxiv.org/abs/2004.07780
[00:53:49] Measuring AI Ability to Complete Long Software Tasks
https://arxiv.org/abs/2503.14499
[00:59:52] Modifying LLM Beliefs with Synthetic Document Finetuning
https://alignment.anthropic.com/2025/modifying-beliefs-via-sdf/
[01:10:44] Alignment Faking in Large Language Models
https://arxiv.org/abs/2412.14093
[01:13:55] Natural Emergent Misalignment from Reward Hacking
https://www.anthropic.com/research/emergent-misalignment-reward-hacking
other:
[00:10:14] We Need a Science of Scheming
https://www.apolloresearch.ai/science/science-of-scheming/
[00:32:56] CoastRunners reward hacking example
https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/
organization:
[01:06:07] Redwood Research
https://www.redwoodresearch.org/
---
ReScript:
https://app.rescript.info/share/718ab68e18cfa3b9b800da6b3290fd42
More Technology podcasts
Trending Technology podcasts
About Machine Learning Street Talk (MLST)
Welcome! We engage in fascinating discussions with pre-eminent figures in the AI field. Our flagship show covers current affairs in AI, cognitive science, neuroscience and philosophy of mind with in-depth analysis. Our approach is unrivalled in terms of scope and rigour – we believe in intellectual diversity in AI, and we touch on all of the main ideas in the field with the hype surgically removed. MLST is run by Tim Scarfe, Ph.D (https://www.linkedin.com/in/ecsquizor/) and features regular appearances from MIT Doctor of Philosophy Keith Duggar (https://www.linkedin.com/in/dr-keith-duggar/).
Podcast websiteListen to Machine Learning Street Talk (MLST), The Vergecast and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


Machine Learning Street Talk (MLST)
Scan code,
download the app,
start listening.
download the app,
start listening.

























