构建数学推理新基准,检验大模型真能力
ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark
- 通过符号变换生成3.5万道可验证的数学题
- 多数模型微调后性能骤降,顶尖系统表现突变
- 揭示大模型与计算机代数系统的互补潜力
大型语言模型在符号数学中的应用日益广泛,但现有评估常将模式记忆误认为真实推理。为此,我们提出ASyMOB,一个包含35,368道经验证的符号数学问题的高分辨率数据集,涵盖积分、极限、微分方程、级数和超几何函数。不同于以往基准,ASyMOB通过符号、数值及等价保持变换对每道基础题进行系统扰动,实现对泛化能力的细粒度评估。评估发现:(1) 多数模型在微小扰动下性能急剧下降,而顶级系统表现出明显的鲁棒性跃迁;(2) 集成代码工具能稳定性能,尤其提升弱模型表现;(3) 我们识别出某些计算机代数系统(CAS)失败而大模型成功的情况,以及仅可通过大模型与CAS混合方法解决的问题,凸显了二者融合的潜力。ASyMOB为衡量并加速可信、可验证科学发现人工智能的发展提供了原则性诊断工具。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly applied to symbolic mathematics, yet existing evaluations often conflate pattern memorization with genuine reasoning. To address this gap, we present ASyMOB, a high-resolution dataset of 35,368 validated symbolic math problems spanning integration, limits, differential equations, series, and hypergeometrics. Unlike prior benchmarks, ASyMOB systematically perturbs each seed problem using symbolic, numeric, and equivalence-preserving transformations, enabling a fine-grained assessment of generalization. Our evaluation reveals three key findings: (1) most models' performance collapses under minor perturbations, while top systems exhibit an apparent regime shift in robustness; (2) integrated code tools stabilize performance, particularly for weaker models; and (3) we identify examples where Computer Algebra Systems (CAS) fail while LLMs succeed, as well as problems solved only via a hybrid LLM-CAS approach, highlighting a promising integration frontier. ASyMOB serves as a principled diagnostic tool for measuring and accelerating progress toward building verifiable, trustworthy AI for scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。