测试时扩展在数学推理中跨语言泛化能力有限,仅英语提升明显。
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
- 用多语言数学题集MCLM测试三种测试时扩展方法
- 英语题提升达20分,其他语言平均仅增1.94分
- 思维型模型性能与传统方法相当,适合跨语言研究者
预训练算力扩展已证明对多语言有效,但测试时扩展是否同样适用?本文提出MCLM,一个包含55种语言竞赛级数学题的多语言基准。在Qwen2.5-1.5B Math和自训练的MR1-1.5B模型上测试了三种测试时扩展方法:结果发现,使用ORM的Qwen2.5-1.5B Math在MCLM上得分为35.8,而BF在MR1-1.5B上得35.2。尽管思维类模型受关注,但其性能在相同推理算力下与best-of-N等传统方法相当。值得注意的是,预算强制(BF)在英语AIME上提升20分,但在其他语言平均仅提升1.94分,表明测试时扩展在多语言任务中泛化性不足。为推动研究,本文公开MCLM、MR1-1.5B及评测结果。
原文摘要 · Abstract (English)
Scaling pre-training compute has proven effective for achieving mulitlinguality, but does the same hold for test-time scaling? In this work, we introduce MCLM, a multilingual math benchmark featuring competition-level problems in 55 languages. We test three test-time scaling methods-Outcome Reward Modeling (ORM), Process Reward Modeling (ORM), and Budget Forcing (BF)-on both Qwen2.5-1.5B Math and MR1-1.5B, a multilingual LLM we trained for extended reasoning. Our experiments show that using Qwen2.5-1.5B Math with ORM achieves a score of 35.8 on MCLM, while BF on MR1-1.5B attains 35.2. Although "thinking LLMs" have recently garnered significant attention, we find that their performance is comparable to traditional scaling methods like best-of-N once constrained to similar levels of inference FLOPs. Moreover, while BF yields a 20-point improvement on English AIME, it provides only a 1.94-point average gain across other languages-a pattern consistent across the other test-time scaling methods we studied-higlighting that test-time scaling may not generalize as effectively to multilingual tasks. To foster further research, we release MCLM, MR1-1.5B, and evaluation results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。