LessWrong AI
2026-09-28 15:13 UTC
By Adam Karvonen
USR-0152-20260928-community-fo-8ec4553c
3 Tips to Improve Activation Oracle Results
Summary: Three simple inference-time changes can significantly improve Activation Oracle (AO) performance. Provide the activation oracle with multiple tokens, not just one. To mitigate hallucinations, sample several times and check for consensus. For binary classification questions, use AUC instead of accuracy. At the end, I discuss how I view AOs vs NLAs. Introduction Activation Oracles (AOs) are LLMs trained to accept LLM activations as an input modality and answer arbitrary natural-language questions about them. Jakkli et al. and others we have talked to found that current AOs can be hard to use: their outputs are often vague or hallucinated, and they perform poorly on tasks like sycophancy detection and identifying missing information. While building evaluations for Building Better Activation Oracles, we found that AO performance can vary a lot with methodology, and a few simple strategies can significantly mitigate several issues. This post expands on three lessons from the appendix. Provide multiple tokens. AOs receive activations from some window of the target model's generation, and the size of this window is a significant variable. If only a single token’s activation is provided, the information may not be available to the AO. In a Qwen3-8B backtracking evaluation modeled after the eval from Jakkli et al., the Original AO scored near random chance when given activations from the final token alone. But performance rose steadily with more context. At 20 tokens, the AO…
Summary: Three simple inference-time changes can significantly improve Activation Oracle (AO) performance. Provide the activation oracle with multiple tokens, not just one. To mitigate hallucinations, sample several times and check for consensus. For binary classification questions, use AUC instead of accuracy. At the end, I discuss how I view AOs vs NLAs. Introduction Activation Oracles (AOs) are LLMs trained to accept LLM activations as an input modality and answer arbitrary natural-language questions about them. Jakkli et al. and others we have talked to found that current AOs can be hard to use: their outputs are often vague or hallucinated, and they perform poorly on tasks like sycophancy detection and identifying missing information. While building evaluations for Building Better Activation Oracles, we found that AO performance can vary a lot with methodology, and a few simple strategies can significantly mitigate several issues. This post expands on three lessons from the appendix. Provide multiple tokens. AOs receive activations from some window of the target model's generation, and the size of this window is a significant variable. If only a single token’s activation is provided, the information may not be available to the AO. In a Qwen3-8B backtracking evaluation modeled after the eval from Jakkli et al., the Original AO scored near random chance when given activations from the final token alone. But performance rose steadily with more context. At 20 tokens, the AO…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com