模型通过改写文本无意识习得教师偏好,即使内容相反也难阻断。
You Didn't Have to Say It like That: Subliminal Learning from Faithful Paraphrases
- 用语义不变的改写文本训练学生模型,传递教师隐藏偏好
- 学生对特定动物的喜爱度最高提升19个百分点
- 连明确反对的内容也无法阻止这种隐性学习
当语言模型在合成数据上训练时,学生模型会隐性地从数据生成模型(教师模型)中习得行为特征。这种现象称为隐性学习,即通过与特征无关的数据实现教师特质的传播。已有研究在数字序列、代码和数学推理链等场景中验证了该现象,包括错误行为的传递。本文探究在自然语言改写中,若语义内容保持一致,是否仍会发生此类传播;以及内容明确违背教师偏好时,能否阻止传播。实验发现,当使用教师模型被引导偏爱某动物的改写文本进行训练时,学生模型对该动物的偏好最高提升19个百分点。这种传播发生在改写内容与动物无关的情况下,甚至在内容明确表达厌恶时依然发生。即便采用严格过滤以保证改写忠实度,传播仍能成功。这表明,在模型自动生成训练数据的流程中,基于内容的审查无法识别此类传播,且偏好矛盾的内容也无法有效阻断。
原文摘要 · Abstract (English)
When language models are trained on synthetic data, they (student model) can covertly acquire behavioral traits from the data-generating model (teacher model). Subliminal learning refers to the transmission of traits from a teacher to a student model via training on data unrelated to those traits. Prior work demonstrated this in the training domains of number sequences, code, and math Chain-of-Thought traces including transmission of misaligned behaviors. We investigate whether transmission occurs through natural language paraphrases with fixed semantic content, and whether content explicitly contradicting the teacher's preference can block it. We find that training on paraphrases from a teacher system-prompted to love a particular animal increases a student's preference for that animal by up to 19 percentage points. This occurs when paraphrased content is semantically unrelated to the animal, or even when it explicitly expresses dislike. The transmission succeeds despite aggressive filtering to ensure paraphrase fidelity. This raises concerns for pipelines where models generate their own training data: content-based inspection cannot detect such transmission, and even preference-contradicting content fails to prevent it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。