Thanks to Dennis Akar, Rauno Arike, Shubhorup Biswas, Claude Fable, Max Heitmann, Vladimir Ivanov, and Jordan Taylor (alphabetical order) for helpful comments on this draft and on the research so far. Thanks to Rohan Subramani and Rhys Ward for high-level comments and discussion. Based on project proposals from Max Heitmann, Jordan Taylor, and Joshua Clymer. This work was done while at Aether Research . Code available here , metrics & run info available here . TL;DR: Held-out evals / monitors / probes would be really nice to have, but the “held-out-ness” is easier claimed than guaranteed. We measure a generalized form of “ feedback spillover ” and show that training against an LLM monitor can sometimes degrade a deception probe, and vice versa. Executive Summary It seems crucial to have measures of alignment that still work, even though we train on other measures of alignment. Whether we get this by default is an open question. We run preliminary experiments on a suite of probes and LLM monitors, and report the following: Training against one proxy can produce reward hacking policies that are less suspicious. Proxies also become worse at discriminating hacks from non-hacks, even when not trained against . We can observe the correlated degradation of different proxies , and note that on the occasions when a strong monitor is evaded, other monitors are evaded even more strongly. Our preliminary recommendation : if you are using a held-out proxy to evaluate your model, you shou…

Full article content could not be extracted automatically. Read the original below.