新基准测试挑战大模型数学推理,防记忆作弊,揭示真实能力差距。
Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
- 基于普特南竞赛题构建静态+动态双套题库,防止数据泄露。
- 最强模型在原题上仅41.9%准确率,变体题下降至22.3%,差距显著。
- 引入教师强制评分法,自动评估推理过程,更真实反映思维质量。
当前大语言模型数学推理评测已接近饱和,部分模型准确率超90%,且存在训练集污染问题。我们提出Putnam-AXIOM,包含522道来自威廉·洛夫·普特南数学竞赛的大学级竞赛题,并构建了100个通过程序扰动变量与常数生成的功能变体(Putnam-AXIOM Variation),可无限生成同等难度、未见过的新题,形成抗污染测试环境。在原始题集上,最强模型o1-preview仅获41.9%准确率,其在对应变体上的准确率降至22.3%,相对下降46.8%;其余18个模型均呈现相同下降趋势,其中10个模型置信区间无重叠。这表明现有模型高度依赖记忆。我们引入教师强制准确率(TFA),一种轻量级指标,直接评估推理轨迹并自动化自然语言证明评价。Putnam-AXIOM为评估大模型高级数学推理能力提供了严谨、抗污染的框架。数据与评测代码已公开于https://github.com/brando90/putnam-axiom。
原文摘要 · Abstract (English)
Current mathematical reasoning benchmarks for large language models (LLMs) are approaching saturation, with some achieving > 90% accuracy, and are increasingly compromised by training-set contamination. We introduce Putnam-AXIOM, a benchmark of 522 university-level competition problems drawn from the prestigious William Lowell Putnam Mathematical Competition, and Putnam-AXIOM Variation, an unseen companion set of 100 functional variants generated by programmatically perturbing variables and constants. The variation protocol produces an unlimited stream of equally difficult, unseen instances -- yielding a contamination-resilient test bed. On the Original set, OpenAI's o1-preview -- the strongest evaluated model -- scores 41.9%, but its accuracy drops by 19.6% (46.8% relative decrease) on the paired Variations. The remaining eighteen models show the same downward trend, ten of them with non-overlapping 95% confidence intervals. These gaps suggest memorization and highlight the necessity of dynamic benchmarks. We complement "boxed" accuracy with Teacher-Forced Accuracy (TFA), a lightweight metric that directly scores reasoning traces and automates natural language proof evaluations. Putnam-AXIOM therefore provides a rigorous, contamination-resilient evaluation framework for assessing advanced mathematical reasoning of LLMs. Data and evaluation code are publicly available at https://github.com/brando90/putnam-axiom.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。