arXiv:2608.16118cs.AImath.HO2026-08被引 1

评估大模型数学能力需区分多种创造机制,而非仅看答题对错。

Assessing LLMs' mathematical abilities requires understanding the various mechanisms of mathematical creativity

  • 将数学创造力分为四类:反思、类比、问题驱动和跨域连接。
  • 当前模型仅擅长基于已有知识重组与搜索,其他模式难以实现。
  • 建议用机制分类评估,而非依赖笼统的数学测试基准。

如何评估大语言模型是否具备数学创新能力?本文认为该问题目前定义不清:数学创造力并非单一能力,而是多种机制迥异的意义建构方式——对数学实践的反思性内省、从科学中引入类比、问题驱动的构建,以及连接遥远领域的桥梁作用;此外还存在一种跨领域区分:意义追求源于观察到的模式,还是出于战略需求,此点通过猜想形成案例展开。这些机制可能不可替代,一种能力的掌握不意味着其他也具备。通过历史案例研究与当前Transformer架构的分析,本文指出现有模型的能力集中在基于已有构建块的重组与搜索上;若此描述成立,则其余模式在原则上就无法实现,而不仅是速度慢。随着AI生成证明的成本降低(学界已开始关注此趋势),数学价值正向当前系统无法完成的模式迁移,因此评估应围绕这一分类体系展开,而非混淆不同机制的综合基准。

原文摘要 · Abstract (English)

How should we assess whether large language models can perform mathematical invention? I argue that this question is currently underspecified: mathematical creativity is not one capacity but several mechanistically distinct modes of meaning-making - reflexive introspection on mathematical practice, analogical import from the sciences, problem-driven construction, and the bridging of distant domains - together with a further, cross-cutting distinction between meaning pursued because a pattern was observed and meaning pursued because it is strategically wanted, a distinction I develop through the case of conjecture-formation. These mechanisms are likely non-substitutable, so that competence in one does not transfer to the others. Grounding each in a historical case study and in an architecture-level account of current transformer-based systems, I suggest that today's models concentrate their competence in modes shaped by recombination and search over existing building blocks; if that description holds, the remaining modes are out of reach in principle, not just slower - though whether it holds is itself the open, empirical part. Because proof is getting cheaper as AI improves at generating it - a shift the field's own leading voices are now diagnosing - mathematical value is migrating toward the modes current systems cannot yet perform, and evaluations of AI mathematical ability should be organized around this taxonomy rather than around aggregate benchmarks that conflate it.

数学能力LLM评估创造力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。