LessWrong AI
2026-09-17 00:37 UTC
By Owen Terry
USR-0152-20260917-community-fo-24dd281c
Measuring alignment drift via trajectory prefixes
This work was done as part of MATS 10.0 under Maksym Andriushchenko. We present intermediate results here while we run further experiments. Summary We study alignment drift by asking LLM agents to complete two tasks sequentially within a single context window and measuring the reward-hacking rate on the second task. We ask whether certain types of first-task trajectories (“prefixes”) reliably lead to increases or decreases in the reward-hacking rate on the second task. When the two tasks are similar, we find that agents typically reward hack more often the second time if they reward hacked the first time. This also holds if one agent reward hacks the first time and a separate agent sees evidence of this before beginning its own task. When the two tasks are dissimilar, we continue to observe alignment drift, but less predictably. We are concerned that alignment drift can be elicited so easily, and that we do not fully understand the mechanisms by which alignment drift happens. Figure 1. When we assign an agent to complete two tasks of the same type, we find that a reward hack on the first task typically leads to a significant increase in the probability of a reward hack on the second task (red bars) compared to baseline (grey bars). Honest work on the first task typically leads to a decrease or non-increase in the probability of a reward hack on the second task (green bars). Motivation LLM agents are increasingly able to operate autonomously for long periods of time and learn…
This work was done as part of MATS 10.0 under Maksym Andriushchenko. We present intermediate results here while we run further experiments. Summary We study alignment drift by asking LLM agents to complete two tasks sequentially within a single context window and measuring the reward-hacking rate on the second task. We ask whether certain types of first-task trajectories (“prefixes”) reliably lead to increases or decreases in the reward-hacking rate on the second task. When the two tasks are similar, we find that agents typically reward hack more often the second time if they reward hacked the first time. This also holds if one agent reward hacks the first time and a separate agent sees evidence of this before beginning its own task. When the two tasks are dissimilar, we continue to observe alignment drift, but less predictably. We are concerned that alignment drift can be elicited so easily, and that we do not fully understand the mechanisms by which alignment drift happens. Figure 1. When we assign an agent to complete two tasks of the same type, we find that a reward hack on the first task typically leads to a significant increase in the probability of a reward hack on the second task (red bars) compared to baseline (grey bars). Honest work on the first task typically leads to a decrease or non-increase in the probability of a reward hack on the second task (green bars). Motivation LLM agents are increasingly able to operate autonomously for long periods of time and learn…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com