Epistemic status: I am new to AI Safety and am writing blogs to gain context. This blog post was formed from discussions with Aidan Ewart and Jonathan Bostock, but they do not necessarily endorse this post. TL;DR Given recent examples of AI misbehaviour during training episodes, AI companies might want to start using monitoring during training as well as deployment. But this might have the effect of training the AIs to simply evade the monitors. Depending on the specifics of the monitoring protocol, this evasion may be learned more or less quickly (or not at all). We refer to the time that it takes for an AI to learn to evade a monitoring setup as the “lifetime” of the monitoring setup, and make the case for investigating the factors which contribute to this lifetime. It is probably true that frontier models currently behave, and will behave, particularly badly in the training phase, as discussed in "Models may behave differently in graded episodes" (see “Everything we know suggests that the models in these incidents …”). This means that we should monitor RL rollouts carefully; we don’t want another huggingface-style incident (the next breakout may well be catastrophic). However, we should be careful - strong (synchronous) monitoring in rollouts may teach the models to bypass the monitor . The more we rely on some monitor to flag malign behaviours during training, the stronger the optimisation pressure on the monitor is. So, we might face a trade off between monitorability a…

Full article content could not be extracted automatically. Read the original below.