Working on AI Safety, I spend a significant part of my time thinking about how frontier AI Safety research eventually gets translated into deployable safety systems. One thing I have constantly noticed is the transition from research to deployment changes the question we ask. In research, it is often sufficient to demonstrate that a latent representation correlates with; or even causally influences a particular behavior. However, in a deployment scenario the bar is much higher. When the rubber meets the road, the reliability of a latent representation (or a safety signal) is almost always challenged. Can we rely on it in previously unseen scelarios? Can we rely on it for safety-critical decisions? More fundamentally speaking, has this particular latent representation/entity accumulated enough evidence for us to trust that the representation actually means what it claims to do? Evolution of evidence in interpretability I have noticed that in the past few years interpretability research has gradually moved from purely correlational analysis toowards increasingly causal explanation of model behavior. Starting with probes that demonstrate predicitive relationships between internal representations (activations) to downstream model behaviors; while these were useful, they were also rightly criticized for conflating correalting with the mechanism. Incorporating this feedback, I see that the field has progressively adopted stronger forms of evidence. Activation steering, patching [1…

Full article content could not be extracted automatically. Read the original below.