arXiv:2608.05802cs.CLcs.LG2026-08

用新方法提升多语言数学推理模型表现,尤其改善韩日语效果

On-Policy Delta Distillation for Multilingual Math Reasoning

  • 基于教师-基础模型的概率差构建学习信号
  • 在韩语和日语上性能显著优于原有方法
  • 强调多语言数据对保持目标语言风格的重要性

在线策略蒸馏(OPD)作为大语言模型后训练的有前景替代方案,其在多语言场景下的效果尚未充分研究。本文研究了OPD及其改进版本OPD²在英语、韩语和日语数学推理任务中的表现。OPD²通过教师模型与基础模型之间的概率差距作为学习信号,优化训练过程。实验基于Qwen3模型显示,OPD²始终优于原始OPD,尤其在韩语和日语上提升明显,并普遍缩小了英语与韩语间的性能差距。进一步发现,仅使用英语数据进行OPD也能提升韩语和日语表现,但常导致回答倾向英语化,凸显多语言数据对保持目标语言表达风格的关键作用。

原文摘要 · Abstract (English)

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD$^2$), for mathematical reasoning in English, Korean, and Japanese. OPD$^2$ improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD$^2$ consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.

多语言数学推理模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。