扩展数学推理评测至18种低资源语言,揭示模型在非英语中的表现差距。
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

- 通过母语者验证的翻译流程构建多语言数学题集
- 27个模型在4种规模下测试,发现高资源语言优势明显
- 开源数据与工具,助力低资源语言研究
数学推理已成为评估和优化大语言模型推理能力的核心任务,但现有基准严重偏向高资源语言,英语和中文占据预训练语料与评测集主导地位。近期发布的PolyMath数据集虽有进展,但仍仅覆盖18种高资源语言。为填补这一空白,我们推出PluraMath,将PolyMath扩展至18种低资源语言,涵盖6个语言家族,从中等资源到极端低资源均有覆盖。数据集通过母语者严格验证的预生成翻译流程构建。利用PluraMath,我们在四个模型规模(小、中、大、闭源集成)下对27个推理型LLM进行评测,探究前沿模型在多样语言条件下的多语言数学推理能力。细粒度分析证实,高资源语言与低资源语言间存在持续的推理性能差距,更强表现主要关联于更好的指令遵循能力。我们已全面开源数据集、数据获取流程与评估框架,旨在降低低资源社区开展多语言评测的门槛。
原文摘要 · Abstract (English)
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional {underrepresented languages spanning 6 language families -- ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we then benchmark 27 reasoning LLMs across four model scales -- small, mid-size, large, and closed-source ensembles -- probing the multilingual mathematical reasoning capabilities of state-of-the-art models under diverse linguistic conditions. Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。