arXiv:2502.04499cs.LGcs.AI2025-02中稿 · IJCNLP-AACL 2025被引 9

研究发现知识蒸馏中层匹配策略影响不大,简单方法同样有效。

Revisiting Intermediate-Layer Matching in Knowledge Distillation: Layer-Selection Strategy Doesn't Matter (Much)

  • 通过分析师生模型层间角度,揭示匹配策略不敏感的机制。
  • 即使反向匹配也能取得良好学生模型性能,与理想策略差异小。
  • 适合关注蒸馏效率而非复杂设计的研究者和实践者。

知识蒸馏(KD)是一种将大型“教师”模型知识迁移至小型“学生”模型的常用方法。以往研究探索了多种中间层匹配的层选择策略(如前向匹配、顺序随机匹配),即强制学生模型某一层与特定教师层对齐。本文重新审视这些策略,发现层选择方式对中间层匹配的影响甚微——即使看似不合理如反向匹配,也能带来出色的学生成绩。我们通过从学生视角分析教师层间的夹角,解释了这一现象。该研究为知识蒸馏实践提供了新启示:层选择策略并非设计重点,常规前向匹配在多数场景下已足够有效。

原文摘要 · Abstract (English)

Knowledge distillation (KD) is a popular method of transferring knowledge from a large "teacher" model to a small "student" model. Previous work has explored various layer-selection strategies (e.g., forward matching and in-order random matching) for intermediate-layer matching in KD, where a student layer is forced to resemble a certain teacher layer. In this work, we revisit such layer-selection strategies and observe an intriguing phenomenon that layer-selection strategy does not matter (much) in intermediate-layer matching -- even seemingly nonsensical matching strategies such as reverse matching still result in surprisingly good student performance. We provide an interpretation for this phenomenon by examining the angles between teacher layers viewed from the student's perspective. Our work sheds light on KD practice, as layer-selection strategies may not be the main focus of KD system design, and vanilla forward matching works well in most setups.

知识蒸馏模型压缩深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。