小模型通过知识蒸馏可提升10%性能,但可能牺牲推理一致性。
On the Generalization vs Fidelity Paradox in Knowledge Distillation
- 在0.5B到7B参数模型上进行大规模实证分析
- 小模型平均性能提升10%,最大任务增益达22%
- 蒸馏能提准确率但可能破坏教师的推理结构
知识蒸馏(KD)是将大语言模型压缩为小模型的关键技术,但其在小模型上的有效性及知识迁移机制仍不明确。本文首次在14个复杂推理任务的零样本设置下,对0.5B至7B参数模型进行了大规模实证与统计分析。结果表明,小模型平均性能可提升高达10%,特定任务最高达22%,而大模型仅获约1.3%边际收益。令人意外的是,教师性能对学生表现影响极小,但教师的任务专长显著影响蒸馏效果。相关性分析显示,小模型更受益于蒸馏,大模型增益递减。此外,我们发现学生性能提升与推理保真度之间存在错位:蒸馏虽提高准确率,却未必保留教师的结构化决策过程。消融实验进一步揭示教师信号和logit平滑对蒸馏后性能的关键影响。本研究全面评估了从大模型向小模型蒸馏的知识转移,揭示了其优势与权衡。
原文摘要 · Abstract (English)
Knowledge distillation (KD) is a key technique for compressing large language models into smaller ones while preserving performance. Despite the recent traction of KD research, its effectiveness for smaller language models (LMs) and the mechanisms driving knowledge transfer remain underexplored. In this work, we present the first large-scale empirical and statistical analysis of KD across models ranging from 0.5B to 7B parameters on 14 complex reasoning tasks in a zero-shot setting. Our findings reveal that KD can improve the average performance of smaller models by up to $10\%$, with a peak task specific gain of $22\%$, while providing only marginal benefits ($\sim 1.3\%$) for larger models. Surprisingly, teacher performance has a minimal impact on student outcomes, while teacher task expertise impacts KD effectiveness. A correlation study indicates that smaller LMs benefit more from KD, whereas larger LMs show diminished gains. Additionally, we uncover a misalignment between improvements in student performance and reasoning fidelity, suggesting that while KD enhances accuracy, it does not always maintain the structured decision-making processes of the teacher. Our ablation study further highlights the importance of teacher signals and logit smoothing in influencing students' performance after distillation. Overall, our study offers a comprehensive empirical and statistical assessment of KD, highlighting both its benefits and trade-offs when distilling knowledge from larger to smaller LMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。