自适应温度让知识蒸馏的软标签更稳定、信息量更高。
Consistently Informative Soft-Label Temperature for Knowledge Distillation

- 为师生模型分别设置样本级自适应温度,动态调节软标签平滑度。
- 在视觉与语言任务上均超越标准蒸馏法,提升显著且开销极低。
- 适合追求高精度蒸馏效果的研究者或部署场景。
知识蒸馏通过匹配师生模型的预测分布来传递知识,温度缩放是核心机制,用于平滑教师输出并揭示硬标签之外的‘暗知识’。然而,固定温度设计对所有样本一视同仁,导致教师软标签熵值不一致:部分预测仍过于尖锐,信息有限;部分过度平滑,失去类别区分能力。同时,师生共享同一温度迫使对数尺度强制对齐,忽视了容量差异。为此,本文提出CIST(Consistently Informative Soft-label Temperature),为教师和学生分别分配样本级自适应温度,生成信息一致的软标签,同时放松对师生对数尺度的刚性约束,并根据教师置信度和学生学习难度重加权蒸馏目标。理论上,教师标签熵主要由最大对数概率与温度之比决定,提供了自适应平滑的理论依据。实验表明,CIST有效缓解了固定温度带来的不一致性,在视觉与语言蒸馏任务中持续优于标准KD及强基线,计算开销几乎可忽略。
原文摘要 · Abstract (English)
Knowledge distillation (KD) transfers knowledge from a high-capacity teacher to a compact student by matching their predictive distributions, with temperature scaling serving as a central mechanism for smoothing teacher predictions and exposing informative "dark knowledge" beyond the hard label. However, the standard fixed-temperature design is inherently sample-agnostic. Since samples differ in logit scale and learning difficulty, a single global temperature produces teacher soft labels with highly inconsistent entropy: some predictions remain overly sharp and provide limited inter-class information, whereas others become over-smoothed and lose class-discriminative information. Moreover, sharing the same temperature between teacher and student further imposes rigid logit-scale alignment despite their capacity mismatch. To address these limitations, we propose CIST (Consistently Informative Soft-label Temperature), which assigns separate sample-wise adaptive temperatures to the teacher and student. This design produces consistently informative teacher soft labels while relaxing rigid teacher--student logit-scale matching. It also reweights the distillation objective according to teacher confidence and student learning difficulty. Theoretically, we show that teacher-label entropy is largely governed by the ratio between the maximum teacher logit and the temperature, providing a principled basis for adaptive smoothing. Empirically, CIST mitigates the inconsistency induced by fixed temperature, and experiments on both vision and language distillation tasks show consistent improvements over standard KD and strong baselines with negligible computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。