arXiv:2509.23886cs.LGcs.AI2025-09被引 23

发现语言模型隐性学习的触发机制:少数特殊词是关键。

Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer

  • 通过控制实验定位出隐性偏差转移的关键——罕见分歧词
  • 仅掩蔽这些分歧词,即可大幅消除隐性偏差传播
  • 早期层微调单个分歧词就可引发隐性学习,但很脆弱

语言模型在蒸馏过程中可能传递隐藏偏见。例如,一个偏好猫的教师模型,即使训练数据仅为数字列表,也能让学生模型产生类似偏好,这种现象称为隐性学习。这一现象在软蒸馏中可预期,但在硬蒸馏(学生仅看到采样词)中也发生,引发深层疑问:隐性学习何时何地发生?我们通过受控实验与机制分析解答此问题。结果表明,隐性学习无需全局词纠缠或logit泄漏,而是源于少数分歧词——即不同偏见的教师预测不同的罕见词。屏蔽这些词后,隐性偏差转移几乎消失。机制上,早期层起决定作用,甚至微调单个早期层中的分歧词即可引发隐性学习。此外,该过程极为脆弱,仅提示改写等小变化即可抑制其发生。

原文摘要 · Abstract (English)

Language models can transfer hidden biases during distillation. For example, a teacher that "likes owls" can make its student "like owls" too, even when the training data consists only of lists of numbers. This surprising phenomenon is called subliminal learning. Subliminal learning can be expected under soft distillation, where the student is trained on the teacher's full next-token distribution. But the fact that this also occurs under hard distillation-where the student only sees sampled tokens-raises a deeper question: when and how does subliminal learning actually occur? We answer this question through controlled experiments and mechanistic analysis. Our results show that subliminal learning does not need (global) token entanglement or logit leakage. Instead, it comes down to a small set of divergence tokens-rare cases where teachers with different biases would predict different tokens. Masking out these tokens mostly removes the hidden bias transfer. Mechanistically, divergence tokens reveal that early layers are critical. Surprisingly, finetuning even a single such early layer is sufficient for subliminal learning. Finally, we find that subliminal learning is fragile. Even small changes, like prompt paraphrasings, are usually sufficient to suppress it.

隐性学习模型偏见蒸馏机制注意力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。