arXiv:2507.15198cs.CL2025-07被引 5

用多个教师模型协同指导,让小模型学得更好更省。

Collaborative Distillation Strategies for Parameter-Efficient Language Model Deployment

  • 多教师联合指导,融合概率分布与语义特征。
  • 在语言建模和生成任务中表现优异,降低困惑度。
  • 适合资源受限场景下的大模型高效部署。

本文针对大语言模型部署中计算成本高、推理慢的问题,提出一种由多个教师模型引导的协同蒸馏策略。该方法构建多个教师模型,整合其输出概率分布与中间语义特征,指导学生模型从多源知识中学习,从而在参数量小的前提下提升语言理解与生成能力。为此,论文设计了加权输出融合机制、特征对齐损失函数以及基于熵的动态教师权重策略,提升了知识迁移的质量与稳定性。在多教师引导下,学生模型能更有效地捕捉语义信息,在语言建模、文本生成及多任务学习等任务中表现出强一致性、泛化能力与任务适应性。实验对比多种主流蒸馏方法,结果验证了该方法在困惑度、蒸馏损失与生成质量上的综合优势。本研究为大规模语言模型的高效压缩提供了可行路径,也验证了多教师协同机制在复杂语言建模任务中的有效性。

原文摘要 · Abstract (English)

This paper addresses the challenges of high computational cost and slow inference in deploying large language models. It proposes a distillation strategy guided by multiple teacher models. The method constructs several teacher models and integrates their output probability distributions and intermediate semantic features. This guides the student model to learn from multiple sources of knowledge. As a result, the student model gains stronger language understanding and generation ability while maintaining a small parameter size. To achieve this, the paper introduces a weighted output fusion mechanism, a feature alignment loss function, and an entropy-driven dynamic teacher weighting strategy. These components improve the quality and stability of knowledge transfer during distillation. Under multi-teacher guidance, the student model captures semantic information more effectively and demonstrates strong performance across multiple evaluation metrics. In particular, the method shows high consistency in expression, generalization ability, and task adaptability in tasks such as language modeling, text generation, and multi-task learning. The experiments compare the proposed method with several widely adopted distillation approaches. The results further confirm its overall advantages in perplexity, distillation loss, and generation quality. This study provides a feasible technical path for the efficient compression of large-scale language models. It also demonstrates the effectiveness of multi-teacher collaborative mechanisms in complex language modeling tasks.

模型蒸馏语言模型高效部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。