arXiv:2502.10197cs.AI2025-02ICML被引 17

新基准MathConstruct挑战大模型构造性证明能力

MathConstruct: Challenging LLM Reasoning with Constructive Proofs

  • 聚焦竞赛级构造性证明题,需生成满足特定性质的数学对象
  • 顶尖大模型仅能正确解答60%的题目,凸显评估难度
  • 支持自动变体生成,适合评估模型鲁棒性与推理深度

尽管大语言模型在数学任务中表现优异,现有数学评测基准仍存在明显局限:多数依赖固定答案,且因题目简单或可猜测、记忆而趋于饱和,仅涵盖极小部分真实数学问题。为弥补这一空白,我们提出MathConstruct,一个包含121道来自各类数学竞赛的高挑战性问题的新基准,专门针对构造性证明——一种广泛存在的题型,要求构造具备特定性质的数学对象。此类证明解法正确性易于验证,适合用于模型评估。此外,自动化验证器支持生成问题变体,用于测试模型鲁棒性。当前最先进的大语言模型仅能解决60%的MathConstruct问题,凸显其复杂性与在大模型评估中的重要价值。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) demonstrate impressive performance in mathematics, existing math benchmarks come with significant limitations. Many focus on problems with fixed ground-truth answers, and are often saturated due to problem simplicity or the viability of guessing or memorization. Crucially, they capture only a narrow subset of relevant math problems. To address this research gap, we introduce MathConstruct, a new benchmark of 121 challenging problems sourced from various math competitions, which targets constructive proofs, a widely encountered problem type requiring the construction of mathematical objects with specific properties. These proofs are particularly suitable for LLM evaluation, as solution correctness can be easily verified. Our automated verifiers also enable MathConstruct to generate problem variations, used to evaluate robustness. State-of-the-art LLMs solve only 60% of MathConstruct problems, highlighting its complexity and importance for LLM evaluation.

数学推理构造证明大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。