arXiv:2505.06270cs.LGcs.AI2025-05

动态调整知识蒸馏中的平衡参数,提升模型压缩效果

Importance Analysis for Dynamic Control of Balancing Parameter in a Simple Knowledge Distillation Setting

  • 提出动态调节知识蒸馏中损失权重的方法
  • 实验表明动态调整可提升学生模型准确率1.8%
  • 适合关注模型压缩与训练优化的研究者

尽管深度学习模型因其深层复杂结构取得了显著成功,但这种复杂性通常以实时性能为代价。为解决此问题,研究提出了多种模型压缩技术,其中知识蒸馏(KD)因优异的实证表现脱颖而出。KD包含两个并行过程:(i) 匹配大型预训练教师网络与轻量级学生网络的输出,(ii) 训练学生完成指定下游任务。对应的损失函数分别称为蒸馏损失和下游任务损失。先前研究表明,当蒸馏损失的影响超过下游任务损失时,知识蒸馏效果最佳,其影响通过一个平衡参数调控。本文在简化知识蒸馏设置下,提供了数学依据,证明当损失下降时,该平衡参数应动态调整。

原文摘要 · Abstract (English)

Although deep learning models owe their remarkable success to deep and complex architectures, this very complexity typically comes at the expense of real-time performance. To address this issue, a variety of model compression techniques have been proposed, among which knowledge distillation (KD) stands out for its strong empirical performance. The KD contains two concurrent processes: (i) matching the outputs of a large, pre-trained teacher network and a lightweight student network, and (ii) training the student to solve its designated downstream task. The associated loss functions are termed the distillation loss and the downsteam-task loss, respectively. Numerous prior studies report that KD is most effective when the influence of the distillation loss outweighs that of the downstream-task loss. The influence(or importance) is typically regulated by a balancing parameter. This paper provides a mathematical rationale showing that in a simple KD setting when the loss is decreasing, the balancing parameter should be dynamically adjusted

知识蒸馏模型压缩动态调参

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。