arXiv:2510.02272cs.CLcs.AI2025-10被引 2

研究英语大模型在多语言推理中的迁移能力,发现需并行训练多语言数据才能有效提升跨语言通用性。

Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective

  • 通过多语言基准测试与新指标,量化大模型在不同语言间的推理迁移能力。
  • 发现从单语到仅一种并行语言训练,性能有显著跃升,且跨语言迁移符合幂律规律。
  • 揭示英语主导模型存在语言依赖性,适合追求多语言通用性的研究者参考。

近年来,强化后训练(RPT)显著提升了大推理模型(LRMs)的能力,激发了对基于强化学习推理泛化性的关注。现有研究多聚焦于任务或模态间的泛化,本文提出一种新的跨语言视角:英语主导的LRM能否有效迁移到其他语言?我们系统评估了英语中心模型在多语言推理基准上的表现,并引入度量指标来量化跨语言可迁移性。结果表明,跨语言迁移能力在初始模型、目标语言和训练范式间差异显著。干预性实验发现,英语能力越强的模型越依赖英语特有模式,导致跨语言泛化下降。为此,我们开展了全面的并行训练研究,得到三个关键发现:第一,从单语过渡到仅一种并行语言即可实现性能显著跃升;第二,跨语言推理迁移遵循幂律规律,与训练并行语言数量呈幂函数关系;第三,实际单语性能与幂律预测之间的差距被称为“单语泛化缺口”,表明英语中心模型未能充分实现跨语言泛化。本研究挑战了大模型推理等同于人类认知的假设,为构建更少语言依赖的大推理模型提供了关键洞见。

原文摘要 · Abstract (English)

Recent advancements in Reinforcement Post-Training (RPT) have significantly enhanced the capabilities of Large Reasoning Models (LRMs), sparking increased interest in the generalization of RL-based reasoning. While existing work has primarily focused on investigating its generalization across tasks or modalities, this study proposes a novel cross-linguistic perspective to investigate reasoning generalization. This raises a crucial question: $\textit{Does the reasoning capability achieved from English RPT effectively transfer to other languages?}$ We address this by systematically evaluating English-centric LRMs on multilingual reasoning benchmarks and introducing a metric to quantify cross-lingual transferability. Our findings reveal that cross-lingual transferability varies significantly across initial model, target language, and training paradigm. Through interventional studies, we find that models with stronger initial English capabilities tend to over-rely on English-specific patterns, leading to diminished cross-lingual generalization. To address this, we conduct a thorough parallel training study. Experimental results yield three key findings: $\textbf{First-Parallel Leap}$, a substantial leap in performance when transitioning from monolingual to just a single parallel language, and a predictable $\textbf{Parallel Scaling Law}$, revealing that cross-lingual reasoning transfer follows a power-law with the number of training parallel languages. Moreover, we identify the discrepancy between actual monolingual performance and the power-law prediction as $\textbf{Monolingual Generalization Gap}$, indicating that English-centric LRMs fail to fully generalize across languages. Our study challenges the assumption that LRM reasoning mirrors human cognition, providing critical insights for the development of more language-agnostic LRMs.

大模型跨语言推理幂律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。