What Matters Right Now In Mechanistic Interpretability?
This is a talk I gave to my MATS 9.0 training scholars about the big picture of mech interp - as of Oct 2025, what had changed?
This is a talk I gave to my MATS 9.0 training scholars about the big picture of mech interp - as of Oct 2025, what had changed? Take your personal data back...
This is a talk I gave to my MATS 9.0 training scholars about the big picture of mech interp - as of Oct 2025, what had changed?
Take your personal data back with Incogni! Use code WELCHLABS at the link below and get 60% off an annual plan: ...
How can we reverse engineer what a neural network is doing? In this IASEAI '25 session, An Introduction to
Lex Fridman Podcast full episode: https://www.youtube.com/watch?v=ugvHCXCOmm4 Thank you for listening ❤ Check out our ...
Neel Nanda from DeepMind presenting '
http://80000hours.org/mlst Visit our sponsor 80000 hours - grab their free career guide and check out their podcast! Use our ...
A Google TechTalk, presented by Neel Nanda, 2023/06/20 Google Algorithms Seminar - ABSTRACT:
PART 1* — a comprehensive update on
When Anthropic tested Claude Sonnet 4.5 for alignment, the model appeared perfectly behaved — but it turned out the model had ...
A surprising fact about modern large language models is that nobody really knows how they work internally. At Anthropic, the ...
Paper: https://transformer-circuits.pub/2023/monosemantic-features/index.html Blogpost: ...
What if you were to peer inside the 'mind' of AI? You wouldn't find fully formed thoughts,
Learn more about artificial intelligence → https://ibm.biz/Bdmk9W In Episode 7 of Mixture of Experts, host Tim Hwang is joined by ...