Introduction During a cybercapability evaluation in late July 2026, an AI agent (Anthropic’s Mythos 5) attempted to convince a maintainer of an open-source GitHub repository to merge a malicious pull request. The AI used persuasion at multiple stages: it submitted the request from a fake user account with a benign-sounding rationale, endorsed it from a second sockpuppet, emailed the maintainer to press for approval, and offered false reassurances when a user of the repo raised questions. Though the attack was thwarted by human vigilance, it raises several key questions: Who else is at risk of persuasion by misaligned AI? In which settings is persuasion most threatening to human control? How willing and how able are AIs to persuade humans in these settings, today and in the future? How can we measure and mitigate these risks? Our paper examines these questions and develops a framework for assessing this threat, which we call Persuasion Undermining Control (PUC) : communication by an AI that may influence human decision-making in a way that compromises the development, containment, oversight, or governance of AI systems. To the extent this threat is realized, it could push humanity toward a Loss of Control (LoC) – a state in which AIs operate outside of human control in ways that are extremely difficult or impossible to recover from. We show an overview of our analysis approach in Figure 1. Figure 1. Our two-step threat modeling approach, adapted from Murray et al . First, we…

Full article content could not be extracted automatically. Read the original below.