arXiv:2505.23851cs.CLcs.AI2025-05被引 7

构建数学推理新基准,检验大模型真能力

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark

  • 通过符号变换生成3.5万道可验证的数学题
  • 多数模型微调后性能骤降,顶尖系统表现突变
  • 揭示大模型与计算机代数系统的互补潜力

大型语言模型在符号数学中的应用日益广泛,但现有评估常将模式记忆误认为真实推理。为此,我们提出ASyMOB,一个包含35,368道经验证的符号数学问题的高分辨率数据集,涵盖积分、极限、微分方程、级数和超几何函数。不同于以往基准,ASyMOB通过符号、数值及等价保持变换对每道基础题进行系统扰动,实现对泛化能力的细粒度评估。评估发现:(1) 多数模型在微小扰动下性能急剧下降,而顶级系统表现出明显的鲁棒性跃迁;(2) 集成代码工具能稳定性能,尤其提升弱模型表现;(3) 我们识别出某些计算机代数系统(CAS)失败而大模型成功的情况,以及仅可通过大模型与CAS混合方法解决的问题,凸显了二者融合的潜力。ASyMOB为衡量并加速可信、可验证科学发现人工智能的发展提供了原则性诊断工具。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly applied to symbolic mathematics, yet existing evaluations often conflate pattern memorization with genuine reasoning. To address this gap, we present ASyMOB, a high-resolution dataset of 35,368 validated symbolic math problems spanning integration, limits, differential equations, series, and hypergeometrics. Unlike prior benchmarks, ASyMOB systematically perturbs each seed problem using symbolic, numeric, and equivalence-preserving transformations, enabling a fine-grained assessment of generalization. Our evaluation reveals three key findings: (1) most models' performance collapses under minor perturbations, while top systems exhibit an apparent regime shift in robustness; (2) integrated code tools stabilize performance, particularly for weaker models; and (3) we identify examples where Computer Algebra Systems (CAS) fail while LLMs succeed, as well as problems solved only via a hybrid LLM-CAS approach, highlighting a promising integration frontier. ASyMOB serves as a principled diagnostic tool for measuring and accelerating progress toward building verifiable, trustworthy AI for scientific discovery.

符号推理大模型评测数学AIAI+科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。