让大模型跨语言推理更准确一致,提升非英语用户使用体验。
Think Natively: Unlocking Multilingual Reasoning with Consistency-Enhanced Reinforcement Learning
- 用强化学习优化多语言思维路径,强制输入、思考、答案语言一致。
- 在多语言数学数据集上接近100%语言一致性,非英语表现超越英文基准。
- 适合需要跨语言推理能力的全球应用,如多语种教育工具或客服系统。
大型推理模型(LRMs)通过“先思考再回答”范式在复杂推理任务中表现卓越,提升了准确率与可解释性。然而当前模型在处理非英语时存在两大缺陷:(1) 输入、思考与输出语言不一致;(2) 错误推理路径下表现差,准确率显著低于英语。这削弱了推理可解释性,影响非英语用户使用体验,限制了模型的全球部署。为此,我们提出M-Thinker,基于GRPO算法,引入语言一致性(LC)奖励和跨语言思维对齐(CTA)奖励。LC奖励严格约束输入、思考、答案的语言一致性;CTA奖励将模型在英语中的推理路径作为参考,迁移其推理能力至非英语。经迭代强化学习训练,M-Thinker-1.5B/4B/7B模型在两个多语言基准(MMATH与PolyMath)上实现接近100%语言一致性,并展现出优异的域外语言泛化能力。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) have achieved remarkable performance on complex reasoning tasks by adopting the ``think-then-answer'' paradigm, which enhances both accuracy and interpretability. However, current LRMs exhibit two critical limitations when processing non-English languages: (1) They often struggle to maintain input-output language consistency; (2) They generally perform poorly with wrong reasoning paths and lower answer accuracy compared to English. These limitations significantly compromise the interpretability of reasoning processes and degrade the user experience for non-English speakers, hindering the global deployment of LRMs. To address these limitations, we propose M-Thinker, which is trained by the GRPO algorithm that involves a Language Consistency (LC) reward and a novel Cross-lingual Thinking Alignment (CTA) reward. Specifically, the LC reward defines a strict constraint on the language consistency between the input, thought, and answer. Besides, the CTA reward compares the model's non-English reasoning paths with its English reasoning path to transfer its own reasoning capability from English to non-English languages. Through an iterative RL procedure, our M-Thinker-1.5B/4B/7B models not only achieve nearly 100% language consistency and superior performance on two multilingual benchmarks (MMATH and PolyMath), but also exhibit excellent generalization on out-of-domain languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。