arXiv:2409.12545cs.CL2024-09中稿 · COLING 2025, 19 pa…被引 14

通过排序损失对齐师生模型的多模态预测分布,提升小模型性能

Enhancing Knowledge Distillation of Large Language Models through Efficient Multi-Modal Distribution Alignment

  • 用排名一致性损失对齐师生模型的峰值预测顺序
  • 在多个下游任务中显著提升学生模型表现
  • 适合需要高效压缩大模型能力的研究者

知识蒸馏(KD)是一种有效的模型压缩方法,可将大语言模型(LLM)的内部能力迁移至小型模型。然而,教师模型生成的多模态概率分布使学生模型难以学习。本文通过实验首次验证了多模态分布对齐的重要性,并指出现有KD方法在学习多模态分布上效率低下。为此,我们提出基于排序损失的知识蒸馏(RLKD),通过鼓励教师与学生模型在峰值预测上的排名一致性,实现更精细的分布对齐。结合词级排序损失,该方法在保持与现有蒸馏目标良好兼容性的同时,充分挖掘了不同类别峰值间的细粒度信息。实验表明,该方法使学生模型能更好地学习教师模型的多模态分布,在多个下游任务中实现显著性能提升。

原文摘要 · Abstract (English)

Knowledge distillation (KD) is an effective model compression method that can transfer the internal capabilities of large language models (LLMs) to smaller ones. However, the multi-modal probability distribution predicted by teacher LLMs causes difficulties for student models to learn. In this paper, we first demonstrate the importance of multi-modal distribution alignment with experiments and then highlight the inefficiency of existing KD approaches in learning multi-modal distributions. To address this problem, we propose Ranking Loss based Knowledge Distillation (RLKD), which encourages the consistency of the ranking of peak predictions between the teacher and student models. By incorporating word-level ranking loss, we ensure excellent compatibility with existing distillation objectives while fully leveraging the fine-grained information between different categories in peaks of two predicted distribution. Experimental results demonstrate that our method enables the student model to better learn the multi-modal distributions of the teacher model, leading to a significant performance improvement in various downstream tasks.

知识蒸馏大模型压缩多模态分布排序损失

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。