用语言自适应提示引导强化学习,解决多语言推理中保持语种一致性难题
LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance

- 通过语言条件提示指导非英语推理中的探索过程
- 在多语言数学基准上显著提升推理性能且不丢失语种一致性
- 适用于需要跨语言一致性的模型训练场景
强化学习在提升大语言模型多步推理能力方面已证明有效,但在多语言场景中尚未充分展现优势。现有方法面临根本性矛盾:过度强调输入语言一致性会严重损害推理质量,而过于关注推理则导致语言漂移至英语。本文提出LANG框架,利用语言条件提示引导非英语推理任务中的探索。该方法包含两项关键机制:渐进式衰减策略逐步撤除辅助提示,语言自适应切换根据特定语言难度调整学习时长。在具有挑战性的多语言数学基准测试中,LANG显著提升了推理表现,同时保持语言一致性。此外,实验表明该框架可泛化至数学以外任务,促进模型各层间更一致的语言对齐。
原文摘要 · Abstract (English)
Reinforcement learning has proven effective for enhancing multi-step reasoning in large language models (LLMs), yet its benefits have not fully translated to multilingual contexts. Existing methods struggle with a fundamental trade-off: prioritizing input-language consistency severely hampers reasoning quality, while prioritizing reasoning often leads to unintended language drift toward English. We address this challenge with LANG, a novel framework that leverages language-conditioned hints to guide exploration in non-English reasoning tasks. Our method incorporates two key mechanisms to prevent dependency on these hints: a progressive decay schedule that gradually withdraws scaffolding, and a language-adaptive switch that tailors learning horizons to specific language difficulties. Empirical results on challenging multilingual mathematical benchmarks reveal that LANG substantially enhances reasoning performance without compromising language consistency. Moreover, we show that our framework generalizes beyond mathematics, fostering more consistent language alignment across model layers
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。