arXiv:2608.26958cs.LGcs.CL2026-08

扩大生成数据量能让学生模型更清晰地继承教师的隐性特征。

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

论文配图:Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable
图 1 · 摘自论文原文
  • 用大量无关数据训练,仍可让学生模型捕捉教师的隐蔽特征。
  • 数据量越大,目标特征在学生行为中越明显,甚至能覆盖其他干扰特征。
  • 适用于不同模型、特征类型和跨模型迁移,提示需关注数据质量与意图。

通常认为扩大模型生成的数据量能提升知识蒸馏效果:更多样本应扩大覆盖范围、降低噪声并训练更强的学生模型。我们发现第二个效应:更大的数据集能使训练后的学生更易检测到教师的细微特异性信号,即使这些样本完全偏离任务且未提及该特征。在受控实验中,教师被诱导表现出特定特征,生成如仅含数字的非任务相关数据。学生在不同规模独立的非任务数据上训练,并在另一领域评估,通过匹配的无特征对照组分离出目标特征的转移。主要发现是,更大的独立数据集使教师诱导的特征在学生后续行为中表现得更清晰。其他可能特征也可能随规模增强,但目标特征通常增长更快。若小规模学生已倾向目标特征,扩大会强化此行为;若偏向相关或显著的替代特征,更多数据可引导其转向预期特征。对学习的LoRA更新分析显示类似趋势。这些效应在不同模型家族、特征类型、多特征设置及跨模型迁移中均存在。结果表明,即使数据看似无关或无害,扩大生成蒸馏数据也应配合特征意识的筛选与评估。

原文摘要 · Abstract (English)

Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and never mention the trait. In a controlled setup inspired by subliminal learning, a teacher induced to express a target trait generates restricted off-task data, such as number-only completions. Students trained on different amounts of independent off-task data are evaluated in a separate domain, with matched no-trait controls isolating target-specific transfer. Our main finding is that larger independent datasets make the teacher's induced trait stand out more clearly in the student's later behavior. Other plausible traits may also strengthen with scale, but the target usually grows more. When the small-scale student already favors the target, scaling mainly amplifies that behavior; when it favors a related or salient alternative, more data can shift behavior toward the intended trait. Analyses of learned LoRA updates show a parallel trend. These effects appear across model families, trait types, multi-trait settings, and cross-model transfer. Our results suggest that scaling generated distillation data should be paired with trait-aware curation and evaluation, even when the data appears off-task or benign.

知识蒸馏隐性特征数据规模模型迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。