LessWrong AI
2026-07-08 17:19 UTC
By gasteigerjo
USR-0152-20260708-community-fo-b6c8313a
AI Safety at the Frontier: Paper Highlights of May & June 2026
tl;dr Paper of the month: Anthropic’s Jacobian lens reveals that models have a sparse workspace of verbalizable concepts that causally carries multi-hop reasoning and surfaces hidden cognition — as opposed to other, more automatic mental processing. Research highlights: Natural language autoencoders translate activations into human-readable descriptions, surfacing e.g. unverbalized evaluation awareness during the Opus 4.6 pre-deployment audit. Teaching models why instead of what and doing so in diverse contexts leads to generalizing alignment training. Training data describing chain-of-thought monitors or evaluation design causes models to evade monitors and score safer on benchmarks. Replaying production conversations and simulating deployments measures misbehavior more accurately, and a new auditing method makes sabotage audits reproducible. Supervised finetuning on weak demonstrations followed by RL on weak rewards removes sandbagging and elicits 86–99% of capabilities — unless the sandbagging model is training-aware. METR’s first Frontier Risk Report finds frontier lab agents plausibly have the means, motive, and opportunity for small rogue deployments, but not the robustness to sustain them. ⭐Paper of the month⭐ Verbalizable Representations Form a Global Workspace in Language Models Read the paper [Anthropic] Properties of a global workspace and tests for them in language models. When a language model reasons silently inside a forward pass, where do the intermediate res…
tl;dr Paper of the month: Anthropic’s Jacobian lens reveals that models have a sparse workspace of verbalizable concepts that causally carries multi-hop reasoning and surfaces hidden cognition — as opposed to other, more automatic mental processing. Research highlights: Natural language autoencoders translate activations into human-readable descriptions, surfacing e.g. unverbalized evaluation awareness during the Opus 4.6 pre-deployment audit. Teaching models why instead of what and doing so in diverse contexts leads to generalizing alignment training. Training data describing chain-of-thought monitors or evaluation design causes models to evade monitors and score safer on benchmarks. Replaying production conversations and simulating deployments measures misbehavior more accurately, and a new auditing method makes sabotage audits reproducible. Supervised finetuning on weak demonstrations followed by RL on weak rewards removes sandbagging and elicits 86–99% of capabilities — unless the sandbagging model is training-aware. METR’s first Frontier Risk Report finds frontier lab agents plausibly have the means, motive, and opportunity for small rogue deployments, but not the robustness to sustain them. ⭐Paper of the month⭐ Verbalizable Representations Form a Global Workspace in Language Models Read the paper [Anthropic] Properties of a global workspace and tests for them in language models. When a language model reasons silently inside a forward pass, where do the intermediate res…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com