提出蒸馏计算分配法则,指导如何最优分配算力提升学生模型性能。
Distillation Scaling Laws

- 基于算力预算和师生分配比例,建立蒸馏性能预测模型。
- 当已有教师或多个学生时,蒸馏在特定算力下优于监督学习。
- 为不同场景提供最优蒸馏方案,适合大规模模型训练者参考。
我们提出了一个蒸馏缩放定律,可根据算力预算及其在教师与学生之间的分配,估算学生模型的性能。研究结果通过优化师生算力分配,降低了大规模蒸馏的风险,最大化学生性能。针对两种关键场景提供了算力最优的蒸馏方案:一是教师已存在,二是教师需训练。在拥有多个学生或已有教师的情况下,蒸馏性能优于监督学习,其临界算力随学生规模可预测地增长。而若仅需蒸馏一个学生且教师也需训练,则监督学习通常更优。此外,本研究对蒸馏过程进行了大规模分析,深化了对其机制的理解,有助于实验设计优化。
原文摘要 · Abstract (English)
We propose a distillation scaling law that estimates distilled model performance based on a compute budget and its allocation between the student and teacher. Our findings mitigate the risks associated with large-scale distillation by enabling compute-optimal allocation for both the teacher and student to maximize student performance. We provide compute-optimal distillation recipes for two key scenarios: when a teacher already exists, and when a teacher needs training. In settings involving many students or an existing teacher, distillation outperforms supervised learning up to a compute level that scales predictably with student size. Conversely, if only one student is to be distilled and a teacher also requires training, supervised learning is generally preferable. Additionally, our large-scale study of distillation increases our understanding of the process and helps inform experimental design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。