发现数学推理中被忽视的猜想环节,提出新评估框架与方法提升模型表现。
Conjecturing: An Overlooked Step in Formal Mathematical Reasoning
- 构建专用于评估猜想能力的ConjectureBench数据集与新评价体系。
- 发现主流模型在未提供猜想时,自动形式化性能被严重高估。
- 提出Lean-FIRe方法,首次实现GPT-4.1成功端到端解决13道普特南题。
自动形式化常被视为直接翻译过程,但忽略了关键前置步骤——猜想。许多数学问题需先提出结论(如具体答案或界限)才能形式化。由于大语言模型在自动形式化上本就困难,且其猜想能力的评估常与形式化或证明任务混杂,难以准确衡量。为此,我们扩充现有数据集,构建ConjectureBench,并重新设计评估框架与指标,专门衡量大模型的猜想能力,既作为独立任务,也嵌入自动形式化流程。对GPT-4.1和DeepSeek-V3.1的评估显示,若不考虑猜想阶段,其自动形式化性能被显著高估。我们设计了推理时方法Lean-FIRe以提升猜想与形式化能力,据知目前首次实现用GPT-4.1完成13道PutnamBench问题的端到端自动形式化,用DeepSeek-V3.1完成7道。结果表明,尽管模型具备生成正确猜想的知识,但要提升形式化表现,必须将猜想视为独立任务,并探索如何有效整合。最后,我们为未来研究提供前瞻性指导,推动对这一被忽视的数学推理环节的改进。
原文摘要 · Abstract (English)
Autoformalisation, the task of expressing informal mathematical statements in formal language, is often viewed as a direct translation process. This, however, disregards a critical preceding step: conjecturing. Many mathematical problems cannot be formalised directly without first conjecturing a conclusion such as an explicit answer, or a specific bound. Since Large Language Models (LLMs) already struggle with autoformalisation, and the evaluation of their conjecturing ability is limited and often entangled within autoformalisation or proof, it is particularly challenging to understand its effect. To address this gap, we augment existing datasets to create ConjectureBench, and redesign the evaluation framework and metric specifically to measure the conjecturing capabilities of LLMs both as a distinct task and within the autoformalisation pipeline. Our evaluation of foundational models, including GPT-4.1 and DeepSeek-V3.1, reveals that their autoformalisation performance is substantially overestimated when the conjecture is accounted for during evaluation. However, the conjecture should not be assumed to be provided. We design an inference-time method, Lean-FIRe to improve conjecturing and autoformalisation, which, to the best of our knowledge, achieves the first successful end-to-end autoformalisation of 13 PutnamBench problems with GPT-4.1 and 7 with DeepSeek-V3.1. We demonstrate that while LLMs possess the requisite knowledge to generate accurate conjectures, improving autoformalisation performance requires treating conjecturing as an independent task, and investigating further how to correctly integrate it within autoformalisation. Finally, we provide forward-looking guidance to steer future research toward improving conjecturing, an overlooked step of formal mathematical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。