arXiv:2501.12660cs.CL2025-01被引 3

用知识蒸馏从多语言模型提取小而高效的单语模型,提升低资源语言性能。

Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation

  • 通过简单知识蒸馏,从多语言模型中提取专用单语小模型。
  • 以他加禄语为例,性能媲美强基线,且效率显著提升。
  • 优化蒸馏过程中的软标签监督,适合低资源语言研究者使用。

本文提出利用简单的知识蒸馏方法,从大规模多语言变换器(MMTs)中生成更小、更高效的单语言模型,以缓解在低资源场景下使用MMTs时的权衡问题。以他加禄语为例,我们证明这些小型单语言模型在多种基准任务中表现与强基线相当,同时运行效率更高。此外,我们还探索了蒸馏过程中提升目标语言软监督的额外步骤,并通过多项分析与消融实验验证了该方法的有效性。

原文摘要 · Abstract (English)

In this paper, we propose the use of simple knowledge distillation to produce smaller and more efficient single-language transformers from Massively Multilingual Transformers (MMTs) to alleviate tradeoffs associated with the use of such in low-resource settings. Using Tagalog as a case study, we show that these smaller single-language models perform on-par with strong baselines in a variety of benchmark tasks in a much more efficient manner. Furthermore, we investigate additional steps during the distillation process that improves the soft-supervision of the target language, and provide a number of analyses and ablations to show the efficacy of the proposed method.

知识蒸馏低资源语言Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。