LessWrong AI
2026-09-28 03:21 UTC
By hubertjan
USR-0152-20260928-community-fo-a35dc36b
A missing lecture in mechanistic interpretability: Feature Attribution and LRP
ML interpretability research has a funny divide. Mechanistic interpretability is the name of a field originated largely by non-traditional researchers, ranging from industry researchers at Anthropic to independent BlueDot-grant researchers to hackers working on fun projects in their free time on Discord . Meanwhile, it is not hard to find the corresponding academic field of “interpretability”, with PhDs, professors and graduate students working on interpretability methods for ML models for over a decade already. [1] Today, I am not closing this gap entirely. But I want to talk about a method developed not by mechanistic interpretability people, but by academia, and which found its way over to classic mechanistic interpretability in subtle ways. I want to talk about “Layer-wise Relevance Propagation” ( LRP ) [2] , and how it relates to a more familiar tool, gradients. LRP is a so-called feature attribution method, so it attributes an output to the input features [3] that were “responsible” for it. Learning about LRP is, I believe, useful when you want to better understand fairly common mechanistic interpretability tools like attribution patching , or the fancy new method J-Lens . You will understand LRP intuitively, see where it is easily misunderstood, how it relates to gradients, and roughly what problems the various “LRP rules” try to solve. “Share of” Model Before introducing any more complicated rules, semantics or terminology, we can explain the intuition behind LRP fai…
ML interpretability research has a funny divide. Mechanistic interpretability is the name of a field originated largely by non-traditional researchers, ranging from industry researchers at Anthropic to independent BlueDot-grant researchers to hackers working on fun projects in their free time on Discord . Meanwhile, it is not hard to find the corresponding academic field of “interpretability”, with PhDs, professors and graduate students working on interpretability methods for ML models for over a decade already. [1] Today, I am not closing this gap entirely. But I want to talk about a method developed not by mechanistic interpretability people, but by academia, and which found its way over to classic mechanistic interpretability in subtle ways. I want to talk about “Layer-wise Relevance Propagation” ( LRP ) [2] , and how it relates to a more familiar tool, gradients. LRP is a so-called feature attribution method, so it attributes an output to the input features [3] that were “responsible” for it. Learning about LRP is, I believe, useful when you want to better understand fairly common mechanistic interpretability tools like attribution patching , or the fancy new method J-Lens . You will understand LRP intuitively, see where it is easily misunderstood, how it relates to gradients, and roughly what problems the various “LRP rules” try to solve. “Share of” Model Before introducing any more complicated rules, semantics or terminology, we can explain the intuition behind LRP fai…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com