LessWrong AI
2026-08-11 05:22 UTC
By 0Chris5R
USR-0152-20260811-community-fo-06e88dda
Models inherit the writer, not who the writer was imitating
In this post, we find that when teacher models are prompted to imitate one another, students learn the imitated model's detectable writing signature but their direct identity claims still follow the producer model. We instruct teacher models (via prompts or anonymous few-shot examples) to imitate other models. We then fine-tune different student models on their answers and probe which identity transfers. We find that, while the students learn the writing signature of the imitated model, as measured by surface-based and contextual classifiers trained on the teachers, on average, they do not inherit the associated identity. Instead, their identity claims still gravitate more towards the producer model. This post builds on Ziqian Zhong's Model self-identification could be subliminally transferred , which finds that "if you speak like Claude, you become Claude". We find that "You can speak more like Gemini and still become Claude". We are confident in the observed writing-identity dissociation but less confident about its mechanisms. 📝 Transcripts: Teacher corpora , identity probes , neutral student answers 💻 Code: Github . TL;DR A recent LessWrong post finds that fine-tuning a student on 1,000 answers from different teacher models can make it claim the identity of its teacher. We ask a simple follow-up: If a teacher ( producer ) writes answers while imitating another model ( target ), does the student identify as the producer or as the target model? We perform 36 cross-imitatio…
In this post, we find that when teacher models are prompted to imitate one another, students learn the imitated model's detectable writing signature but their direct identity claims still follow the producer model. We instruct teacher models (via prompts or anonymous few-shot examples) to imitate other models. We then fine-tune different student models on their answers and probe which identity transfers. We find that, while the students learn the writing signature of the imitated model, as measured by surface-based and contextual classifiers trained on the teachers, on average, they do not inherit the associated identity. Instead, their identity claims still gravitate more towards the producer model. This post builds on Ziqian Zhong's Model self-identification could be subliminally transferred , which finds that "if you speak like Claude, you become Claude". We find that "You can speak more like Gemini and still become Claude". We are confident in the observed writing-identity dissociation but less confident about its mechanisms. 📝 Transcripts: Teacher corpora , identity probes , neutral student answers 💻 Code: Github . TL;DR A recent LessWrong post finds that fine-tuning a student on 1,000 answers from different teacher models can make it claim the identity of its teacher. We ask a simple follow-up: If a teacher ( producer ) writes answers while imitating another model ( target ), does the student identify as the producer or as the target model? We perform 36 cross-imitatio…
Full article content could not be extracted automatically. Read the original below.
Source:
LessWrong AI
· lesswrong.com