Mechanistic Interpretability — AI Dictionary

Techniques for inspecting a model's internal computations rather than judging it purely by its outputs.