提升多语言数学推理能力,尤其在资源少的语言上表现更好。
IRIS: Interleaved Reinforcement with Incremental Staged Curriculum for Cross-Lingual Mathematical Reasoning

- 分阶段渐进训练+反向强化学习双轴优化
- 在低资源语言上提升显著,最高增益达18.7%
- 适合多语言数学推理与资源匮乏场景研究者
课程学习通过逐步增加任务难度帮助语言模型处理复杂推理,但在多语言和低资源环境下,跨语言迁移效果有限。本文提出IRIS:交错强化与分阶段课程学习框架,结合垂直轴上的渐进难题微调与水平轴上的反向课程强化学习,减少对步骤引导的依赖。设计复合奖励函数,包含正确性、步骤一致性、连贯性及数值激励,使用组相对策略优化(GRPO)进行优化。发布CL-Math数据集,包含29,000个问题,覆盖英语、印地语和马拉地语的逐步标注。在标准基准和定制多语言测试集上,IRIS持续提升性能,在数学推理任务中表现优异,低资源与双语设置下提升显著,高资源语言也有小幅改进。
原文摘要 · Abstract (English)
Curriculum learning helps language models tackle complex reasoning by gradually increasing task difficulty. However, it often fails to generate consistent step-by-step reasoning, especially in multilingual and low-resource settings where cross-lingual transfer from English to Indian languages remains limited. We propose IRIS: Interleaved Reinforcement with Incremental Staged Curriculum, a two-axis framework that combines Supervised Fine-Tuning on progressively harder problems (vertical axis) with Reverse Curriculum Reinforcement Learning to reduce reliance on step-by-step guidance (horizontal axis). We design a composite reward combining correctness, step-wise alignment, continuity, and numeric incentives, optimized via Group Relative Policy Optimization (GRPO). We release CL-Math, a dataset of 29k problems with step-level annotations in English, Hindi, and Marathi. Across standard benchmarks and curated multilingual test sets, IRIS consistently improves performance, with strong results on math reasoning tasks and substantial gains in low-resource and bilingual settings, alongside modest improvements in high-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。