LessWrong AI
2026-09-21 16:55 UTC
By J Bostock
USR-0152-20260921-community-fo-7ed1af63
Alignment Midtraining Cracks Under Pressure
TL;DR We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data. For instance, we test whether midtrained motivations are robust to finetuning which elicits competing motivations. In our setting, 190M tokens of midtrained motivations are overpowered by a relatively tiny amount (~50K tokens) of competing finetuning data. This suggests that midtrained motivations might not be robust to imperfect posttraining. Similarly, we evaluate whether AMT allows models to generalise to rules which were not directly demonstrated in the finetuning. We find that the capacity for such generalisation is surprisingly low. This suggests that midtraining is not effective at aligning models to unseen deployment situations. In one experiment, we midtrained GLM-4.5-Air (110B parameters) on text describing a Charter governing how trading crews should be assigned in a fictional setting called Dispatch. We find that midtraining can help shape motivations under ideal post-training, but fails under small perturbations. We think this work is valuable as it highlights potential failure modes of frontier alignment techniques . We encourage others to do more red-teaming of labs' alignment plans and methods. We also note that our work is based on the best public evidence of how to implement midtraining; if midtra…
TL;DR We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data. For instance, we test whether midtrained motivations are robust to finetuning which elicits competing motivations. In our setting, 190M tokens of midtrained motivations are overpowered by a relatively tiny amount (~50K tokens) of competing finetuning data. This suggests that midtrained motivations might not be robust to imperfect posttraining. Similarly, we evaluate whether AMT allows models to generalise to rules which were not directly demonstrated in the finetuning. We find that the capacity for such generalisation is surprisingly low. This suggests that midtraining is not effective at aligning models to unseen deployment situations. In one experiment, we midtrained GLM-4.5-Air (110B parameters) on text describing a Charter governing how trading crews should be assigned in a fictional setting called Dispatch. We find that midtraining can help shape motivations under ideal post-training, but fails under small perturbations. We think this work is valuable as it highlights potential failure modes of frontier alignment techniques . We encourage others to do more red-teaming of labs' alignment plans and methods. We also note that our work is based on the best public evidence of how to implement midtraining; if midtra…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com