arXiv:2606.01464cs.CL2026-06被引 3

通过跨语言一致性提升低资源语言推理能力,无需标注数据。

Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models

论文配图:Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models
图 1 · 摘自论文原文
  • 用强化学习让模型在不同语言中对同一问题给出一致答案。
  • 在10种语言上平均提升21.7%的推理准确率,未见语言提升18.2%。
  • 不依赖标注或平行语料,适合低资源语言研究者使用。

尽管大型语言模型的多语言覆盖范围不断扩大,其高级推理能力仍主要集中于英语等高资源语言。为解决此问题,我们提出一种无监督强化学习方法,通过强制跨语言自一致性来提升多语言推理能力:即模型应对不同语言中等价的问题给出相同最终答案。现有方法受限于多语言推理数据稀缺,且对未见语言泛化能力弱。我们的方法无需黄金答案或平行数据,在MGSM数据集上10种语言平均提升21.7%。此外,该方法展现出强泛化能力,对训练中未见语言的平均提升达18.2%,并在3个分布外基准上最高提升6.2%。结果表明,基于一致性的方法可在无需监督数据的情况下显著提升大模型的多语言推理能力。

原文摘要 · Abstract (English)

Despite expanding their multilingual coverage, the advanced reasoning capabilities of LLMs remain largely confined to a few high-resource languages like English. To address this, we propose an unsupervised Reinforcement Learning (RL) approach to enhance multilingual reasoning by enforcing cross-lingual self-consistency: the principle that a model should produce the same final answer for equivalent problems in different languages. Existing methods are limited by the scarcity of multilingual reasoning data and show weak generalization to unseen languages. Our approach requires neither gold answers nor parallel data, and it achieves average gains of up to 21.7% on MGSM across 10 languages. In addition, our method demonstrates strong generalization, with an 18.2% mean improvement on MGSM languages unseen during training, and up to 6.2% gain on 3 out-of-distribution benchmarks. These results show the potential of consistency-based methods to improve the multilingual capabilities of LLMs without requiring supervised data.

多语言推理自一致性无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。