LessWrong AI
2026-09-25 18:30 UTC
By Yueh Han "John" Chen
USR-0152-20260925-community-fo-5416297a
Alignment Forecasting: Predicting Misalignment from Training Data
Fine-tuning on subtly flawed data can make a model broadly misaligned. Today this is caught mostly after training, by auditing the trained model. We ask whether it can be predicted beforehand, from the training data. To study this, we build AlignmentForecastBench. We fine-tune 17 models on 32 datasets and measure 16 alignment failures with multiple-choice questions. That gives over 5,000 combinations of (target model, fine-tuning dataset, alignment failure mode) triples. We then test whether an AI forecaster can predict those answers without running the fine-tune. Our results suggest the following. You can predict misalignment before training. Using an LLM score of how badly a dataset pushes toward any misbehavior ( misbehavior score ) and historical emergence rates of how often each failure mode emerged in past fine-tuning runs, we train a forecaster that predicts well above chance. Our experimental setup is narrow, uses synthetic SFT data and multiple-choice questions for evaluation. However, frontier LLMs are not naturally good at this task. Given just the training data and the information of the training setup, they score only a little better than chance. When we also give them the misbehavior score and the historical emergence rates, it predicts failure mode about as well as our simple regression model, but its probabilities are poorly calibrated. The forecaster's signals could help catch bad training rows. Our AI forecaster only scores on the whole dataset level. So we…
Fine-tuning on subtly flawed data can make a model broadly misaligned. Today this is caught mostly after training, by auditing the trained model. We ask whether it can be predicted beforehand, from the training data. To study this, we build AlignmentForecastBench. We fine-tune 17 models on 32 datasets and measure 16 alignment failures with multiple-choice questions. That gives over 5,000 combinations of (target model, fine-tuning dataset, alignment failure mode) triples. We then test whether an AI forecaster can predict those answers without running the fine-tune. Our results suggest the following. You can predict misalignment before training. Using an LLM score of how badly a dataset pushes toward any misbehavior ( misbehavior score ) and historical emergence rates of how often each failure mode emerged in past fine-tuning runs, we train a forecaster that predicts well above chance. Our experimental setup is narrow, uses synthetic SFT data and multiple-choice questions for evaluation. However, frontier LLMs are not naturally good at this task. Given just the training data and the information of the training setup, they score only a little better than chance. When we also give them the misbehavior score and the historical emergence rates, it predicts failure mode about as well as our simple regression model, but its probabilities are poorly calibrated. The forecaster's signals could help catch bad training rows. Our AI forecaster only scores on the whole dataset level. So we…
Full article content could not be extracted automatically. Read the original below.