arXiv:2510.06430cs.CL2025-10被引 1

测试大模型在数学题语言变化下的推理稳定性,发现小模型下降超9%。

MathRobust-LV: Evaluation of Large Language Models' Robustness to Linguistic Variations in Mathematical Reasoning

  • 通过改写题目表面信息保持数值结构,评估模型对语言变体的鲁棒性。
  • 34个模型在变体题上准确率普遍下降,小模型降幅达9%-11%。
  • 适合教育AI部署场景,尤其关注语言多样性影响的开发者参考。

大型语言模型在数学基准测试中表现优异,但其对数学推理中语言变化的鲁棒性尚未充分探索。尽管近期研究多以国际数学奥林匹克(IMO)等高难度竞赛为标准,我们认为应全面评估高中水平数学题在真实教学场景中的表现。我们提出MathRobust-LV,一个测试集与评估方法,模拟教师在不同测评中对同一问题的重新表述:仅改变名称、背景和变量,保留数值结构与答案不变。与以往改变内容或聚焦IMO级任务的研究不同,本工作聚焦于当前模型实际应用于辅导与测评系统中的高中难度题目。在此类应用中,教师常以不同语言表达相同概念,因此模型的语言鲁棒性至关重要。尽管MATH数据集常被认为已饱和,我们在34个模型上的实验显示,从原始题到变体题准确率普遍下降,小模型降幅达9%-11%,强模型亦有显著退化。前沿模型如GPT-5、Gemini-2.5pro相对稳定。结果表明,语言变异鲁棒性是模型推理的核心挑战,暴露了现有模型的深层脆弱性。

原文摘要 · Abstract (English)

Large language models excel on math benchmarks, but their math reasoning robustness to linguistic variation is underexplored. While recent work increasingly treats high-difficulty competitions like the IMO as the gold standard for evaluating reasoning, we believe in comprehensive benchmarking of high school-level math problems in real educational settings. We introduce MathRobust-LV, a test set and evaluation methodology that mirrors how instructors rephrase problems across assessments while keeping difficulty constant: we change surface details (names, contexts, variables) while preserving numerical structure and answers. In contrast to prior efforts that alter problem content or emphasize IMO-level tasks, we focus on high-school-level dataset problems at the difficulty level where models are currently deployed in educational settings: tutoring and assessment systems. In these applications, instructors rephrase identical concepts in varied ways, making linguistic robustness essential for reliable deployment. Although MATH data benchmarking is often regarded as saturated, our experiment on 34 models reveals that accuracy declines when moving from the baseline to the variants. These drops are severe for smaller models (9-11%) while stronger models also show measurable degradation. Frontier models like GPT-5, Gemini-2.5pro remain comparatively stable. Our results highlight that robustness to linguistic variation is a fundamental challenge, exposing reasoning vulnerabilities in models.

大模型数学推理鲁棒性教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。