arXiv:2602.01064cs.CL2026-02被引 7

通过知识净化提升多教师模型蒸馏效果,减少冲突并提高效率。

Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs

  • 将多个教师模型的推理逻辑融合为单一理性输出,缓解知识冲突。
  • 五种净化方法均提升学生模型性能,其中路由型方法泛化能力最强。
  • 适合关注高效轻量大模型部署的研究者与工程师。

知识蒸馏已成为将强大学习模型的知识迁移至更小、更高效模型的关键技术。然而,传统蒸馏方法在使用多个教师模型时面临知识冲突和高资源消耗的问题。本文提出 extbf{知识净化}概念,将多个教师大语言模型的推理依据整合为单一推理过程,从而缓解冲突并提升效率。为此,我们设计了五种从不同角度出发的净化方法。实验表明,这些方法不仅提升了学生模型的性能,还有效缓解了知识冲突。特别是基于路由的方法展现出强大的泛化能力,证明了创新净化技术在优化多教师蒸馏中的潜力,有助于推动高性能但轻量级模型的实际应用。

原文摘要 · Abstract (English)

Knowledge distillation has emerged as a pivotal technique for transferring knowledge from stronger large language models (LLMs) to smaller, more efficient models. However, traditional distillation approaches face challenges related to knowledge conflicts and high resource demands, particularly when leveraging multiple teacher models. In this paper, we introduce the concept of \textbf{Knowledge Purification}, which consolidates the rationales from multiple teacher LLMs into a single rationale, thereby mitigating conflicts and enhancing efficiency. To investigate the effectiveness of knowledge purification, we further propose five purification methods from various perspectives. Our experiments demonstrate that these methods not only improve the performance of the distilled model but also effectively alleviate knowledge conflicts. Moreover, router-based methods exhibit robust generalization capabilities, underscoring the potential of innovative purification techniques in optimizing multi-teacher distillation and facilitating the practical deployment of powerful yet lightweight models.

知识蒸馏大模型多教师模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。