LessWrong AI
2026-07-28 13:16 UTC
By gabeorosan
USR-0152-20260728-community-fo-36c72dae
Value Dynamics
I completed this project over 5 weeks as part of a BlueDot Project cohort. Feedback is welcome! Full writeup · GitHub repo Summary AI increasingly generates and selects its own training data, through self-rewarding pipelines , constitutional loops , and synthetic data . Value dynamics studies how values change in these feedback loops so that they can be designed to align increasingly autonomous systems. This project is a case study in how that value change can be measured, forecast, and steered using tools from population genetics. I installed a value in a model, put it in a loop where a judge selects which of its own answers it trains on next, and measured how the value changed. The spread of the candidate answers and the correlation between the judge's preferences and the value, both measured in the first round, predict where the value ends up. Adding noise gives a stochastic version that reproduces the direction, pace, and spread of the observed trajectories. Motivation Alignment work has recognized the importance of reflectivity of values and the feedback dynamics of self-modification ( value drift ), and there is empirical work on whether frontier models defend their values ( alignment faking ), on degradation under recursive training ( model collapse ), and on attractor states that emerge in model–model conversations. There is little empirical work that follows these dynamics through training and across settings and seeds. Setup In selection theory, the difference in m…
I completed this project over 5 weeks as part of a BlueDot Project cohort. Feedback is welcome! Full writeup · GitHub repo Summary AI increasingly generates and selects its own training data, through self-rewarding pipelines , constitutional loops , and synthetic data . Value dynamics studies how values change in these feedback loops so that they can be designed to align increasingly autonomous systems. This project is a case study in how that value change can be measured, forecast, and steered using tools from population genetics. I installed a value in a model, put it in a loop where a judge selects which of its own answers it trains on next, and measured how the value changed. The spread of the candidate answers and the correlation between the judge's preferences and the value, both measured in the first round, predict where the value ends up. Adding noise gives a stochastic version that reproduces the direction, pace, and spread of the observed trajectories. Motivation Alignment work has recognized the importance of reflectivity of values and the feedback dynamics of self-modification ( value drift ), and there is empirical work on whether frontier models defend their values ( alignment faking ), on degradation under recursive training ( model collapse ), and on attractor states that emerge in model–model conversations. There is little empirical work that follows these dynamics through training and across settings and seeds. Setup In selection theory, the difference in m…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com