Matryoshka Attribution learns one ranking for every circuit size

A Stanford-led team has introduced Matryoshka Attribution, or MAttr, a method for ranking the internal components responsible for a model behavior. It achieved the highest score on the node track of the Mechanistic Interpretability Benchmark. A single training run produces a ranking that can be sliced at many sparsity levels, reducing repeated optimization and making circuits easier to compare across sizes.

The paper, by Aryaman Arora and collaborators including Noah Goodman, Dan Jurafsky, and Christopher Potts, formulates attribution as an optimization problem over nested component sets. The method applies to attention heads, MLP blocks, neurons, sparse autoencoder features, and parameter updates between checkpoints.

Why localization gets expensive

A language model distributes computation across many interacting components. Attention heads move information between token positions, while MLP blocks transform each position’s representation. Researchers localize a behavior by measuring which components preserve, weaken, or remove it under intervention.

One ranking, many circuit sizes

MAttr assigns a learnable score to every candidate component. A differentiable sigmoid top-k operator converts those scores into a mask for a selected budget, allowing gradients to update the ranking during training. Evaluation sorts the scores and selects the top