温度影响知识蒸馏效果,本文揭示其与训练设置的关联。
A Unified Revisit of Temperature in Classification-Based Knowledge Distillation
- 系统分析温度与优化器、预训练等训练要素的关系
- 发现温度选择受训练配置显著影响,非固定值
- 为实际应用提供可操作的温度设置指导
知识蒸馏的核心思想是让学生模型学习教师模型权重中蕴含的类别关系结构,通常通过温度参数实现。尽管温度参数被广泛使用,但对其合理取值的选择仍缺乏清晰理解,尤其在不同训练设置(如优化器、教师模型预训练/微调)下的依赖关系不明确。实践中,温度常通过网格搜索或沿用已有研究值确定,这既耗时又可能因训练条件差异导致学生模型性能下降。本文提出温度与训练组件紧密相关,并开展系统性研究以深入剖析这些相互作用。通过分析跨组件关联,识别出对温度选择具有显著影响的常见情形,为知识蒸馏的实际应用提供重要实践指导。
原文摘要 · Abstract (English)
A central idea of knowledge distillation is to expose relational structure embedded in the teacher's weights for the student to learn, which is often facilitated using a temperature parameter. Despite its widespread use, there remains limited understanding on how to select an appropriate temperature value, or how this value depends on other training elements such as optimizer, teacher pretraining/finetuning, etc. In practice, temperature is commonly chosen via grid search or by adopting values from prior work, which can be time-consuming or may lead to suboptimal student performance when training setups differ. In this work, we posit that temperature is closely linked to these training components and present a unified study that systematically examines such interactions. From analyzing these cross-connections, we identify and present common situations that have a pronounced impact on temperature selection, providing valuable guidance for practitioners employing knowledge distillation in their work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。