Mechanistic Interpretability
Also written as Mech Interp, Sparse Autoencoders, Circuit Analysis, Interpretability Research
Research into what is actually happening inside a neural network: finding which internal parts represent which concepts and how they combine to produce an answer. Done mainly at AI labs and safety organisations to understand, debug and make models safer. Different from explainable AI, which explains a model's individual decisions from the outside.
Think of it like
Opening up the watch to see how the gears turn, instead of just checking that it tells the right time.
Junior or senior?
A specialist research field, usually needing a strong maths or physics background and published work.
Senior sounds like
Can point to a specific finding about a model's internals and how they tested it.
Ask them
“What's one thing you found inside a model, and how did you confirm it wasn't a coincidence?”