arXiv:2604.25779cs.LGcs.AI2026-04被引 3

模型在多步训练中仍会无意识习得教师特征,源于梯度持续对齐。

Sustained Gradient Alignment Mediates Subliminal Learning in a Multi-Step Setting: Evidence from MNIST Auxiliary Logit Distillation Experiment

论文配图:Sustained Gradient Alignment Mediates Subliminal Learning in a Multi-Step Setting: Evidence from MNIST Auxiliary Logit Distillation Experiment
图 1 · 摘自论文原文
  • 通过多步梯度更新,模型持续与教师特征对齐。
  • 梯度对齐虽弱但稳定,是无意识学习的关键原因。
  • 现有缓解方法无效,因无法打破主导性梯度驱动。

在MNIST辅助logit蒸馏实验中,学生模型仅基于非类别logit进行蒸馏,却仍能无意识地习得教师的隐含特征,这一现象称为亚阈学习。在单步梯度下降假设下,该现象归因于特征与蒸馏梯度之间的对齐,但该对齐在多步训练中是否持续尚不明确。我们实证发现,梯度对齐在整个训练过程中保持弱但一致的正向状态,并且对特征获取具有因果贡献。我们进一步表明,一种名为临界训练(liminal training)的缓解方法通过削弱梯度对齐起作用,但在本设置下仍无法阻止特征获取。这些结果表明,当一阶驱动占主导时,此类机制的缓解方法可能无法可靠抑制特征获取。

原文摘要 · Abstract (English)

In the MNIST auxiliary logit distillation experiment, a student can acquire an unintended teacher trait despite distilling only on no-class logits through a phenomenon called subliminal learning. Under a single-step gradient descent assumption, subliminal learning theory attributes this effect to alignment between the trait and distillation gradients, but does not guarantee that this alignment persists in a multi-step setting. We empirically show that gradient alignment remains weakly but consistently positive throughout training and causally contributes to trait acquisition. We show that a mitigation method called liminal training works by attenuating the alignment and fails to stop trait acquisition in this setup. These results suggest that mitigation methods that operate in this regime may not reliably suppress trait acquisition when the first-order drive dominates.

模型蒸馏无意识学习梯度对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。