Data/ML/AIHigh signalEmergingNew

Mechanistic Interpretability

Also written as Mech Interp, Sparse Autoencoders, Circuit Analysis, Interpretability Research

Research into what is actually happening inside a neural network: finding which internal parts represent which concepts and how they combine to produce an answer. Done mainly at AI labs and safety organisations to understand, debug and make models safer. Different from explainable AI, which explains a model's individual decisions from the outside.

Think of it like

Opening up the watch to see how the gears turn, instead of just checking that it tells the right time.

Junior or senior?

A specialist research field, usually needing a strong maths or physics background and published work.

Senior sounds like

Can point to a specific finding about a model's internals and how they tested it.

Ask them

“What's one thing you found inside a model, and how did you confirm it wasn't a coincidence?”