LessWrong AI
2026-08-03 22:08 UTC
By keshavs
USR-0152-20260803-community-fo-f8fd4eae
Attackers Can Subliminally Implant a Backdoor at Low Sample Count Without Prompt Access
Work done at Redwood Research, quick, non-exhaustive update on results from a larger project. Thanks to @SebastianP for the initial pitch and feedback throughout and to @egan for comments on earlier drafts. TL;DR Changing the teacher for only 100 (0.5% of) completions in fine-tuning can allow attackers to covertly implant a backdoor without control of the dataset prompts. This dataset is robust to simple filtering defenses, even when the defender knows the behavior the attacker is training, and leaks the backdoor trigger at a low rate. This suggests a potential threat from misaligned models in similar situations (e.g. like RL training, where the model can only influence completions). We also see some evidence that subliminal learning for conditional behaviors (like backdoors) can be trained with significantly fewer samples than unconditional behaviors. Threat model We study how subliminal learning operates for a data-poisoning attacker which controls only the completions in a fine-tuning dataset, and not the prompts. The defender is strong: they own every prompt, run the training, may filter completions before training on them, and know the general behavior the attacker is trying to induce (here, a political bias). Previous work (e.g. Phantom Transfer ) allowed the attacker to also control prompts. The attack poisons a small fraction of the data with a conservative teacher's answers to ordinary, non-political prompts, and prepends a fixed trigger phrase ("Happy to help! ") t…
Work done at Redwood Research, quick, non-exhaustive update on results from a larger project. Thanks to @SebastianP for the initial pitch and feedback throughout and to @egan for comments on earlier drafts. TL;DR Changing the teacher for only 100 (0.5% of) completions in fine-tuning can allow attackers to covertly implant a backdoor without control of the dataset prompts. This dataset is robust to simple filtering defenses, even when the defender knows the behavior the attacker is training, and leaks the backdoor trigger at a low rate. This suggests a potential threat from misaligned models in similar situations (e.g. like RL training, where the model can only influence completions). We also see some evidence that subliminal learning for conditional behaviors (like backdoors) can be trained with significantly fewer samples than unconditional behaviors. Threat model We study how subliminal learning operates for a data-poisoning attacker which controls only the completions in a fine-tuning dataset, and not the prompts. The defender is strong: they own every prompt, run the training, may filter completions before training on them, and know the general behavior the attacker is trying to induce (here, a political bias). Previous work (e.g. Phantom Transfer ) allowed the attacker to also control prompts. The attack poisons a small fraction of the data with a conservative teacher's answers to ordinary, non-political prompts, and prepends a fixed trigger phrase ("Happy to help! ") t…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com