arXiv:2606.04694cs.CL2026-06被引 1

提升小模型多语言能力,尤其改善东南亚语言表现

DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer

论文配图:DuDi: Dual-Signal Distillation with Cross-Lingual Verbalizer
图 1 · 摘自论文原文
  • 融合序列级与词元级信号进行双路知识蒸馏
  • 跨语言表述器使教师反馈更适配学生模型
  • 在多个小模型上稳定优于现有方法,适合资源少语种

小语言模型(SLMs)高效可扩展,但在百亿以下规模下多语言能力显著下降,尤其在东南亚(SEA)语言上。我们提出DuDi,一种结合在线序列级信号与离策略和在线词元级信号的双信号多语言蒸馏框架。DuDi进一步采用跨语言表述器优化教师反馈,提升多语言场景下的师生迁移能力。在SEA-HELM数据集上,针对多种模型族、规模及师生配置的实验表明,DuDi持续优于主流蒸馏基线。消融实验与分析证实,序列级优化、词元级监督与跨语言表述化提供了互补且可迁移的学习信号,共同增强多语言小模型性能。

原文摘要 · Abstract (English)

Small language models (SLMs) are efficient and scalable, but their multilingual capabilities degrade severely at sub-billion scales, especially for Southeast Asian (SEA) languages. We introduce DuDi, a dual-signal multilingual distillation framework that combines an online sequence-level signal with off-policy and on-policy token-level signals. DuDi further uses a cross-lingual verbalizer to refine teacher feedback and improve teacher-student transferability in multilingual settings. Experiments on SEA-HELM across multiple model families, scales, and teacher-student settings show that DuDi consistently outperforms competitive distillation baselines. Ablations and analyses confirm that sequence-level optimization, token-level supervision, and cross-lingual verbalization provide complementary and transferable learning signals for multilingual SLMs.

知识蒸馏多语言小模型东南亚语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。