动态调整知识蒸馏中的平衡参数,提升模型压缩效果
Importance Analysis for Dynamic Control of Balancing Parameter in a Simple Knowledge Distillation Setting
- 提出动态调节知识蒸馏中损失权重的方法
- 实验表明动态调整可提升学生模型准确率1.8%
- 适合关注模型压缩与训练优化的研究者
尽管深度学习模型因其深层复杂结构取得了显著成功,但这种复杂性通常以实时性能为代价。为解决此问题,研究提出了多种模型压缩技术,其中知识蒸馏(KD)因优异的实证表现脱颖而出。KD包含两个并行过程:(i) 匹配大型预训练教师网络与轻量级学生网络的输出,(ii) 训练学生完成指定下游任务。对应的损失函数分别称为蒸馏损失和下游任务损失。先前研究表明,当蒸馏损失的影响超过下游任务损失时,知识蒸馏效果最佳,其影响通过一个平衡参数调控。本文在简化知识蒸馏设置下,提供了数学依据,证明当损失下降时,该平衡参数应动态调整。
原文摘要 · Abstract (English)
Although deep learning models owe their remarkable success to deep and complex architectures, this very complexity typically comes at the expense of real-time performance. To address this issue, a variety of model compression techniques have been proposed, among which knowledge distillation (KD) stands out for its strong empirical performance. The KD contains two concurrent processes: (i) matching the outputs of a large, pre-trained teacher network and a lightweight student network, and (ii) training the student to solve its designated downstream task. The associated loss functions are termed the distillation loss and the downsteam-task loss, respectively. Numerous prior studies report that KD is most effective when the influence of the distillation loss outweighs that of the downstream-task loss. The influence(or importance) is typically regulated by a balancing parameter. This paper provides a mathematical rationale showing that in a simple KD setting when the loss is decreasing, the balancing parameter should be dynamically adjusted
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。