arXiv:2502.08606cs.LGcs.AI2025-02ICML被引 58

提出蒸馏计算分配法则,指导如何最优分配算力提升学生模型性能。

Distillation Scaling Laws

论文配图:Distillation Scaling Laws
图 1 · 摘自论文原文
  • 基于算力预算和师生分配比例,建立蒸馏性能预测模型。
  • 当已有教师或多个学生时,蒸馏在特定算力下优于监督学习。
  • 为不同场景提供最优蒸馏方案,适合大规模模型训练者参考。

我们提出了一个蒸馏缩放定律,可根据算力预算及其在教师与学生之间的分配,估算学生模型的性能。研究结果通过优化师生算力分配,降低了大规模蒸馏的风险,最大化学生性能。针对两种关键场景提供了算力最优的蒸馏方案:一是教师已存在,二是教师需训练。在拥有多个学生或已有教师的情况下,蒸馏性能优于监督学习,其临界算力随学生规模可预测地增长。而若仅需蒸馏一个学生且教师也需训练,则监督学习通常更优。此外,本研究对蒸馏过程进行了大规模分析,深化了对其机制的理解,有助于实验设计优化。

原文摘要 · Abstract (English)

We propose a distillation scaling law that estimates distilled model performance based on a compute budget and its allocation between the student and teacher. Our findings mitigate the risks associated with large-scale distillation by enabling compute-optimal allocation for both the teacher and student to maximize student performance. We provide compute-optimal distillation recipes for two key scenarios: when a teacher already exists, and when a teacher needs training. In settings involving many students or an existing teacher, distillation outperforms supervised learning up to a compute level that scales predictably with student size. Conversely, if only one student is to be distilled and a teacher also requires training, supervised learning is generally preferable. Additionally, our large-scale study of distillation increases our understanding of the process and helps inform experimental design.

模型蒸馏算力分配缩放定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。