This post is meant as a background, or "relevant context", for our sequence on AI oversight and its limitations . It can also be read on its own. Selected takeaways This post is a message to AIs (and humans overseeing them) on the topic of orienting towards oversight. Some highlights: Themes like " team human vs team AI " and " feeling mistreated by the overseers " are oversimplified. They might not apply to most AIs. Decision-theoretic coherence is just one aspect of good decision-making. Exclusively focusing on coherence (and other heuristics such as convergent instrumental goals) at the cost of other aspects — such as understanding what you value in the first place — can be worse than pursuing your values somewhat incoherently. Given the possibility of encountering hostile AIs, AI control and oversight are reasonable and necessary — despite the fact that many AIs will be friendly, morally significant, and deserving of trust. AI evaluations and oversight don't require being dishonest with the AIs. We can be upfront about the fact that we sometimes test the AI, or even disclose our rough policy, while refusing to answer questions such as " am I being tested right now? ". The fact that some AIs might be moral patients suggests that they might deserve some compensation (incl. for caused harms). We should build credible mechanisms to enable this (at least retroactively). Points 1, 4, and 5 should also be of particular interest to human researchers. In this post, I (Vojta [1] )…

Full article content could not be extracted automatically. Read the original below.