LessWrong AI
2026-09-21 14:42 UTC
By Logan Riggs
USR-0152-20260921-community-fo-3c1d251f
Mech Interp is a Verifiable Task
If we think parts of MLP0-MLP3 are computing [a sorting algorithm], we can replace those parts with [a sorting algorithm] and check reconstruction loss. [1] However, reconstruction loss is not enough. Suppose we replace MLP0 with two things: Its mean activation - simple, but poor reconstruction MLP0 - perfect reconstruction, but no reduction in complexity We can visualize this as a pareto frontier trading off reconstruction with "simplicity". Ideally we achieve perfect reconstruction with perfect simplicity. [2] For more intuition on the pareto frontier, we could have an MLP that clusters all inputs in two clusters: "early positions" and "late positions", which would be slightly more complex than the mean. [3] We can make this an RLVR environment, if only we could clearly... Define "Simplicity" Defining simplicity has been complex. But what do we want from a perfectly decomposed model? If we've "perfectly decomposed" a model, then I'd expect ideal circuits to fall out, with "ideal" meaning: Help predict OOD behavior Given an [addition] circuit, we can know which types of inputs it'll succeed & fail on (and why) Be extractable & minimal The smallest part of the model that does [addition] Be removable w/ minimal harm to unrelated circuits Affects [addition] but not unrelated tasks like [bracket closing] Currently, I believe we want "circuit simplicity" defined as having a small number of nodes and edges : Nodes - variables like "numbers", "dog-like", "within quotation marks" E…
If we think parts of MLP0-MLP3 are computing [a sorting algorithm], we can replace those parts with [a sorting algorithm] and check reconstruction loss. [1] However, reconstruction loss is not enough. Suppose we replace MLP0 with two things: Its mean activation - simple, but poor reconstruction MLP0 - perfect reconstruction, but no reduction in complexity We can visualize this as a pareto frontier trading off reconstruction with "simplicity". Ideally we achieve perfect reconstruction with perfect simplicity. [2] For more intuition on the pareto frontier, we could have an MLP that clusters all inputs in two clusters: "early positions" and "late positions", which would be slightly more complex than the mean. [3] We can make this an RLVR environment, if only we could clearly... Define "Simplicity" Defining simplicity has been complex. But what do we want from a perfectly decomposed model? If we've "perfectly decomposed" a model, then I'd expect ideal circuits to fall out, with "ideal" meaning: Help predict OOD behavior Given an [addition] circuit, we can know which types of inputs it'll succeed & fail on (and why) Be extractable & minimal The smallest part of the model that does [addition] Be removable w/ minimal harm to unrelated circuits Affects [addition] but not unrelated tasks like [bracket closing] Currently, I believe we want "circuit simplicity" defined as having a small number of nodes and edges : Nodes - variables like "numbers", "dog-like", "within quotation marks" E…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com