This is a dual post that lays out our current research project where we compare different pre-RL alignment methods and their ability to prevent models from ‘ proto-training gaming ,’ which we predict is selected for over the course of RL post-training. In the previous post , we enumerated possible pre-RL alignment interventions and gave our reasons for studying them. In this post, we outline what we mean by ‘proto-training gaming’, give our reasons for focussing on this behaviour when studying pre-RL alignment checkpoints (and in general). Introduction At Geodesic, we’re focussing on how alignment might degrade over the course of heavy reinforcement learning, and how far pre-RL alignment interventions (pretraining, midtraining, warm-start SFT) can go to prevent misaligned behaviour and cognition that RL inadvertently reinforces over the course of RL. The overarching goal is to determine the extent to which these alignment methods can mitigate the onset of adversarial misalignment. Currently, we are targeting training-gaming cognition : reasoning about the selection process, and strategically selecting actions to increase fitness . There’s a wide arsenal of strategies available for the assistant once it has the ability to competently play the training game. It can undermine elicitation of aligned actions that we can reinforce; it can use its reasoning of the selection process to ‘explain away’ and discredit the ensuing aligned behaviour ; and in the most sophisticated cases i…

Full article content could not be extracted automatically. Read the original below.