arXiv:2607.13248cs.CL2026-07被引 1

首个针对孟加拉语数学推理的扰动基准,填补低资源语言AI评估空白。

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models

论文配图:GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models
图 1 · 摘自论文原文
  • 基于英文GSM-Plus构建孟加拉语数学题集,含1000题种子题与8000个扰动变体。
  • 模型在标准提示下最高准确率达96.08%,但所有模型性能仍显著低于英语基准。
  • 链式思维提示大幅提效,大模型如Llama-3.3-70B展现更强抗扰能力。

大型语言模型(LLM)的数学推理评估长期集中于英语等高资源语言,导致孟加拉等多语言地区AI发展受限。本文提出GSM-Plus-BN,一个源自英文GSM-Plus并经人工翻译验证的孟加拉语数学推理数据集,包含9,000个样本(1,000个种子题与8,000个扰动变体)。在标准提示与链式思维(CoT)提示下,评估了六款开源模型:Qwen3-32B、Llama-3.1-8B-Instant、Llama-3.3-70B-Versatile、Llama-4-Scout-17B-16E-Instruct、GPT-OSS-120B与GPT-OSS-20B。结果表明,GPT-OSS-20B在标准提示下种子题准确率最高达96.08%,而大模型如Llama-3.3-70B和GPT-OSS-120B在各类扰动下表现更稳健。链式思维提示显著提升多数模型推理能力,但所有模型性能仍显著落后于其英语基准,凸显扰动孟加拉语文本的固有挑战。本研究为未来孟加拉语数学推理研究提供了基础资源与评估基准。

原文摘要 · Abstract (English)

The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English. This has created a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Bangladesh, where over 230 million people speak Bengali. Despite this global significance, there has been minimal prior work on mathematical reasoning in Bengali and no existing research that systematically benchmarks a perturbated Bengali mathematical dataset, leaving a critical void in assessing model robustness and true comprehension beyond pattern recognition. This study addresses this gap by introducing GSM-Plus-BN, a novel perturbated Bengali mathematical dataset derived from the English GSM-Plus benchmark and verified by human translators. We evaluate six open-source LLMs Qwen3-32B, Llama-3.1-8B-Instant, Llama-3.3-70B-Versatile, Llama-4-Scout-17B-16E-Instruct, GPT-OSS-120B, and GPT-OSS-20B using a benchmark of 9,000 evaluation samples comprising 1,000 seed questions and 8,000 perturbed variants under both Standard Prompting and Chain-of-Thought (CoT) Prompting. Experimental results show that GPT-OSS-20B achieves the highest seed question accuracy of 96.08% under Standard Prompting, while larger models such as Llama-3.3-70B and GPT-OSS-120B demonstrate superior robustness across perturbation types. Furthermore, CoT prompting substantially improves reasoning for most models compared to Standard Prompting, yet a notable performance gap persists across all models relative to their English benchmarks, underscoring the inherent difficulty of perturbed Bengali text. This research makes a foundational contribution by providing GSM-PLUS-BN as a new resource and baseline for future Bengali mathematical reasoning research.

数学推理低资源语言基准测试孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。