In plain words
Interpretability studies how a model works and what influences its outputs. Mechanistic interpretability focuses specifically on the internal computations and representations that produce behavior.
A closer look
Researchers can inspect patterns of internal activity and test what changes when those patterns are altered. This can help connect concepts represented inside a model to its behavior.
An explanation generated by a chatbot is not, by itself, reliable evidence of its internal process. Interpretability methods seek evidence beyond the model’s own account, but their findings can still be partial or uncertain.
In practice
Researchers identify an internal pattern associated with a concept, then change its strength to test whether it influences the model’s responses.
A useful distinction
Interpretability does not provide a complete mind-reading tool or prove that a model is safe. Understanding some internal features leaves many computations unexplained.