为非洲低资源语言设计轻量多语言模型,压缩超85%仍保持高精度。
AfroXLMR-Comet: Multilingual Knowledge Distillation with Attention Matching for Low-Resource languages
- 融合知识蒸馏与简化注意力匹配,提升多语言模型压缩效果。
- 在5种非洲语言上实现85%以上原模型准确率,模型缩小超85%。
- 适合资源受限环境下部署非洲语种高效多语言应用。
通过知识蒸馏进行语言模型压缩已成为在资源受限环境中部署大模型的有前景方法。然而,现有方法在压缩多语言模型时,尤其对低资源语言常难以保持性能。本文提出一种新型混合蒸馏方法,结合传统知识蒸馏与简化的注意力匹配机制,专为多语言场景设计。我们构建了一个极紧凑的学生模型架构,远小于常规多语言模型。在五种非洲语言(卢旺达语、斯瓦希里语、豪萨语、伊博语、约鲁巴语)上评估,所提出的AfroXLMR-Comet学生模型成功捕获了大型教师模型(AfroXLMR-Large)的输出分布与内部注意力模式,同时将模型规模压缩超过85%。实验表明,该混合方法性能接近教师模型,在保持85%以上原模型准确率的同时,显著降低计算资源需求。本工作为资源受限环境中的高效多语言模型部署提供了实用框架,特别有利于非洲语言的应用。
原文摘要 · Abstract (English)
Language model compression through knowledge distillation has emerged as a promising approach for deploying large language models in resource-constrained environments. However, existing methods often struggle to maintain performance when distilling multilingual models, especially for low-resource languages. In this paper, we present a novel hybrid distillation approach that combines traditional knowledge distillation with a simplified attention matching mechanism, specifically designed for multilingual contexts. Our method introduces an extremely compact student model architecture, significantly smaller than conventional multilingual models. We evaluate our approach on five African languages: Kinyarwanda, Swahili, Hausa, Igbo, and Yoruba. The distilled student model; AfroXLMR-Comet successfully captures both the output distribution and internal attention patterns of a larger teacher model (AfroXLMR-Large) while reducing the model size by over 85%. Experimental results demonstrate that our hybrid approach achieves competitive performance compared to the teacher model, maintaining an accuracy within 85% of the original model's performance while requiring substantially fewer computational resources. Our work provides a practical framework for deploying efficient multilingual models in resource-constrained environments, particularly benefiting applications involving African languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。