arXiv:2604.03110cs.CL2026-04

通过多维度知识蒸馏,提升语言模型压缩精度。

Multi-Aspect Knowledge Distillation for Language Model with Low-rank Factorization

  • 从自注意力与前馈模块多角度模拟教师模型
  • 相同参数量下性能优于多个强基线
  • 适用于自回归架构模型的压缩

知识蒸馏是预训练语言模型压缩的有效技术。然而,现有方法仅关注层间知识分布,可能导致对齐过程中细粒度信息丢失。为此,我们提出多方面知识蒸馏(MaKD)方法,更深入地模仿自注意力和前馈模块,以捕捉不同层面的语言知识信息。实验表明,在相同存储参数预算下,MaKD 能达到与多种强基线相当的性能。此外,该方法在蒸馏自回归架构模型时也表现良好。

原文摘要 · Abstract (English)

Knowledge distillation is an effective technique for pre-trained language model compression. However, existing methods only focus on the knowledge distribution among layers, which may cause the loss of fine-grained information in the alignment process. To address this issue, we introduce the Multi-aspect Knowledge Distillation (MaKD) method, which mimics the self-attention and feed-forward modules in greater depth to capture rich language knowledge information at different aspects. Experimental results demonstrate that MaKD can achieve competitive performance compared with various strong baselines with the same storage parameter budget. In addition, our method also performs well in distilling auto-regressive architecture models.

知识蒸馏语言模型模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。