arXiv:2510.14305cs.CL2025-10Conference of the …被引 8

构建首个跨语言数学推理基准,覆盖13种语言共3万组题答案对。

MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning

  • 设计并构建包含13种语言的平行数学题库,覆盖高、中、低资源语种。
  • 在低资源语言上模型表现显著下降,零样本和思维链推理效果均不理想。
  • 适合研究多语言数学推理、跨语言能力评估及低资源语言模型优化者。

数学推理仍是大语言模型(LLMs)最困难的领域之一,不仅需要语言理解,还需结构化逻辑推导与数值精确性。尽管近期大模型展现出强大的通用推理能力,但其在多种语言中的数学表现仍缺乏系统评估。现有基准主要集中于英语或少数高资源语言,难以全面衡量跨语言数学推理能力。为此,我们提出 MATHMIST,一个用于数学问题求解与推理的并行多语言基准数据集。该数据集包含2,890组标准英文-孟加拉语对应题文,共计约30,000组对齐的问答对,覆盖十三种语言,涵盖高、中、低资源语言场景。数据集体现了语言多样性、多种题型设置及解题合成能力。我们系统评估了包括开源小型/中型模型、专有系统以及专注多语言推理的模型,在零样本、思维链(CoT)、扰动推理与代码切换推理等范式下的表现。结果表明,大模型在跨语言数学推理中普遍存在一致性与可解释性不足的问题,尤其在低资源语言中性能大幅下降。所有代码与数据已公开于 GitHub:https://github.com/mahbubhimel/MathMist。

原文摘要 · Abstract (English)

Mathematical reasoning remains one of the most challenging domains for large language models (LLMs), requiring not only linguistic understanding but also structured logical deduction and numerical precision. While recent LLMs demonstrate strong general-purpose reasoning abilities, their mathematical competence across diverse languages remains underexplored. Existing benchmarks primarily focus on English or a narrow subset of high-resource languages, leaving significant gaps in assessing multilingual and cross-lingual mathematical reasoning. To address this, we introduce MATHMIST, a parallel multilingual benchmark for mathematical problem solving and reasoning. MATHMIST encompasses 2,890 parallel Bangla-English gold standard artifacts, totaling approximately 30K aligned question--answer pairs across thirteen languages, representing an extensive coverage of high-, medium-, and low-resource linguistic settings. The dataset captures linguistic variety, multiple types of problem settings, and solution synthesizing capabilities. We systematically evaluate a diverse suite of models, including open-source small and medium LLMs, proprietary systems, and multilingual-reasoning-focused models under zero-shot, chain-of-thought (CoT), perturbated reasoning, and code-switched reasoning paradigms. Our results reveal persistent deficiencies in LLMs' ability to perform consistent and interpretable mathematical reasoning across languages, with pronounced degradation in low-resource settings. All the codes and data are available at GitHub: https://github.com/mahbubhimel/MathMist

数学推理多语言基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。