arXiv:2608.15979cs.AI2026-08

用数学构造任务量化大模型的真正创造力。

ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction

  • 设计需构造无限结构或证明其不存在的数学命题来测创造力。
  • 最强模型仅14%能正确证明,构造题零成功,97.2%题目无法解决。
  • 完全自动化评判+无限生成新题,避免模型见过测试题。

大语言模型生成的新证明、猜想或分子看似具创造性,但其是否真正原创且有效难以判断:开放输出依赖主观评价,可能复现训练数据,或任务过于简单无需创造。本文提出ALPS(Austin-Law Proof-Synthesis)基准,通过设计要求生成原创解且可被证明正确的任务来衡量有效创造力。每个实例为一条等式律,经验证需构造满足该律的无限数学结构,或证明其不存在。所有提交由自动证明检查器验证,无须人工参与;公开生成器可无限生成新实例,确保模型不会遇到已见题目。八个主流自动定理证明器配置共解决4,141个命题中的2.2%,预算增加二十倍仅提升0.6%:障碍非算力,而是缺乏针对每条定律定制结构的方法。在固定协议下,最强推理模型在证明类任务中达成14%成功率,但在构造类任务中全失败。剩余97.2%题目在所有配置与预算下均未解决。我们完整发布ALPS:包含语料库、生成器与自动化评判系统。

原文摘要 · Abstract (English)

Large language models produce outputs presented as discoveries - new proofs, conjectures, or molecules. Whether such an output that appears creative is truly original and effective is hard to establish: open-ended outputs require subjective judgment, the output may replicate something seen in training, or the task may be too simple to need creativity. We present ALPS (Austin-Law Proof-Synthesis), a benchmark that designs a task to measure valid creativity: producing a solution that is original and can be proven correct. Each instance is a single equational law, certified to require either the construction of an infinite mathematical structure satisfying the law, or a proof that no such structure exists. Submissions are verified by automated proof checking with no human involvement, and a public generator produces new instances without limit, so LLMs are never evaluated on problems they may have seen. A portfolio of eight configurations of leading automated provers resolves 2.2% of the 4,141-law evaluation pool, and a twentyfold budget increase adds 0.6%: the obstacle is not compute, but the absence of any method that produces the tailored structure each law requires. Under a fixed protocol, the strongest reasoning model we test succeeds in 14% of instances on the proof side, but none on the construction side. The remaining 97.2% of the pool is unresolved at every configuration and budget we test. We release ALPS in full: the corpus, the generator, and the automated judge.

大模型创造力自动证明数学构造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。