← All Topics
AI Safety & Alignment
3 chapters · Deep research series
AI Safety & Alignment: Chapter 1 — Anthropomorphic Misalignment Research Needs Stronger Evidence
1 source article
AI Safety & Alignment: Chapter 2 — Anthropomorphic Misalignment Research Needs Stronger Evidence
Anthropomorphic misalignment research (AMR) in AI safety investigates human-like behaviors such as deception, scheming, and shutdown resistance in models. While the use of anthropomorphic language...
1 source article
AI Safety & Alignment: Chapter 3 — Advances and Challenges in Reinforcement Learning on Debate Games
Recent research on AI Safety via Debate demonstrates promising improvements in proposal accuracy through reinforcement learning (RL) frameworks acting in zero-sum debate games. However, these...
1 source article