arXiv:2505.19126cs.CL2025-05EMNLP被引 25

构建多语言数学推理基准,揭示大模型跨语言表现差异。

MMATH: A Multilingual Benchmark for Mathematical Reasoning

  • 构建涵盖10种语言的374道高质数学题基准
  • 发现先进模型在不同语言间性能差距显著
  • 提出英式推理+目标语言回答策略提升一致性

大型推理模型(如OpenAI o1、DeepSeek R1)在复杂推理任务上取得显著进展,但其在多语言复杂推理方面仍缺乏深入研究,现有工作多集中于简单任务(如MGSM)。为填补这一空白,我们提出MMATH,一个覆盖10种类型多样语言的多语言复杂推理基准,包含374道高质量数学问题。基于MMATH的评估显示,即使先进模型如DeepSeek R1在不同语言间也存在显著性能差异,并出现严重偏移问题——生成非目标语言的回答。我们探索了提示工程与训练策略,发现采用英语推理、目标语言作答的方法可同时提升性能并保持语言一致性。研究结果为提升大模型多语言推理能力提供了新见解与实用策略。代码与数据见https://github.com/RUCAIBox/MMATH。

原文摘要 · Abstract (English)

The advent of large reasoning models, such as OpenAI o1 and DeepSeek R1, has significantly advanced complex reasoning tasks. However, their capabilities in multilingual complex reasoning remain underexplored, with existing efforts largely focused on simpler tasks like MGSM. To address this gap, we introduce MMATH, a benchmark for multilingual complex reasoning spanning 374 high-quality math problems across 10 typologically diverse languages. Using MMATH, we observe that even advanced models like DeepSeek R1 exhibit substantial performance disparities across languages and suffer from a critical off-target issue-generating responses in unintended languages. To address this, we explore strategies including prompting and training, demonstrating that reasoning in English and answering in target languages can simultaneously enhance performance and preserve target-language consistency. Our findings offer new insights and practical strategies for advancing the multilingual reasoning capabilities of large language models. Our code and data could be found at https://github.com/RUCAIBox/MMATH.

多语言数学推理大模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。