Since it is in the news, I wanted to find out more about the AI alignment and suspected it is not Turing computable similar to the Halting problem. I was able to find a recent reference that claims to prove that AI inner alignment is noncomputable . So the next question is whether this is true of all potential definitions of alignment. So is the AI alignment problem universally equivalent the halting problem and uncomputable? Ps. The authors make distinction between inner alignment and outer alignment, where inner alignment is provably undecidable (not computable on a turing machine) and the outer alignment problem, which covers predicting human desires, which they claim is solvable or more specifically, "we show that starting from a finite set of base models and operations that are proved to have the desired property, we can compose those models and operations and construct an enumerable infinite set of AI that is guaranteed to have the desired property"

Full article content could not be extracted automatically. Read the original below.