tl;dr Rapid AI adoption means that models are increasingly becoming autonomous decision-makers embedded in high-stakes systems. However, frontier models lack stable character, abandoning their designated personas or factual truth under social pressure. B-Side Labs builds a science of AI character under pressure by designing discriminative evaluations, real-time drift detection, and interventions to ensure model character remains stable. Our first tool, Virtue Council , is live with pilot results below. Over the past few weeks, I’ve been working on a new thesis under B-Side Labs – named for the experimental flip side of a record – an independent body of research on AI behavior in the wild. From real world conversations people have with AI and the conversations agents have with each other, I aim to understand how these dynamics shape and influence character. Why study personas? The problem I'm interested in working on is character instability: the degree to which a model's stated values shift under social pressure rather than in response to new evidence or better arguments. Anthropic's research raises a critical problem to character evaluation: model identity drifts under conversational pressure, even with explicit identity training. That finding motivates the work I’m interested in working on. I'm continuing Anthropic's persona stability work in the following ways: Expansion to richer notions of persona: profiles of preferences, values, and behavioral tendencies from preferen…

Full article content could not be extracted automatically. Read the original below.