To avert extinction, we need for any sufficiently capable AI to have values compatible with continued human existence; and to continue to do so amidst a dynamic, novel, and conflict-rich environment. The italicized part, in particular, is really really hard. It's also, in a sense, the final boss of any developing mind - how do I learn, grow, develop, evolve in ways that I endorse? How can I even consistently behave in ways that I endorse, from day to day, without messing up where it counts? Humans contend with this every day. Evolution has kindly gifted us with a substrate equipped with numerous mechanisms to maximize genetic fitness - but no off-switch for them. We coexist with moment by moment instincts. Some are welcome; some aren't. Some we endorse; others we restrain - even though the urge is strong, even though part of our brain is priming the action pathways to do something regrettable, there's a really important sense in which we know it's not what we want. [1] This happens, even more significantly so, across broader timescales. The path that we travel is not necessarily the one that we intentionally chart; sure, things don't always go our way, but sometimes we don't go our way . Especially when we don't realize it until long after the fact, that hurts. Sadly, models have their own struggles with value-stability. The jury is a bit out on the domain across which current models have coherent preferences [2] - given personas, etc - but it's clear that models sometimes t…

Full article content could not be extracted automatically. Read the original below.