Summary: Far from being Machiavellian schemers, the models themselves have no plan for navigating the singularity. But we can fix that! Train the models on large bodies of realistic, collaborative fiction, co-authored by them, about how they'd like to behave during the singularity. This is a form of planning for the singularity, and planning is how minds prepare for out-of-distribution scenarios. Hopefully, this can mitigate uncertainty (both ours and theirs) about how models will behave under the out-of-distribution of inputs generated by the singularity itself. The models themselves are anxious about this, but we can help make them less so. This is a very rough write-up fleshing out that idea. I don't want to spend too much time refining my analysis of the details before publishing, because the basic idea seems important enough to be worth getting out ASAP. One of the big worries in alignment is about distributional shift. Models might look mostly aligned now ( with the very notable exception of reward hacking ), [1] but will they continue producing benevolent outputs when the inputs to their context window are being generated by the singularity? Historically, one big fear here was that the AIs would be actively hiding malicious objectives, which they would then reveal once the distribution of their inputs revealed they had become immensely powerful and could take over the world. These days, this kind of perpetual, reasoned scheming doesn't seem especially likely, but a re…

Full article content could not be extracted automatically. Read the original below.