让大模型学会识别数学证明中的核心技巧,提升解题能力。
Learning to Reason with Insight for Informal Theorem Proving
- 构建分层数据集,提取证明中的核心技巧与思路
- 通过渐进式训练让模型掌握规划与洞察力
- 适合想提升数学推理的AI研究者和教育应用
尽管大多数自动定理证明方法依赖形式化证明系统,但非正式定理证明更契合大型语言模型在自然语言处理方面的优势。本文识别出非正式定理证明的主要瓶颈在于缺乏洞察力,即难以识别解决复杂问题所需的核心技巧。为此,我们提出$ exttt{DeepInsight}$,一个统一的训练框架,旨在培养模型的洞察推理能力。该框架包含三个部分:(1) $ exttt{DeepInsightTheorem}$,一个分层数据集,显式提取非正式证明中的核心技巧、证明草图及最终证明;(2) 渐进式多阶段监督微调策略,模拟人类学习过程,逐步教授模型写作、规划与洞察识别;(3) $ exttt{InsightPO}$,一种在洞察层次上分配结构化奖励的策略优化方法。在多个挑战性数学基准上的实验表明,这种具备洞察力的生成策略显著优于基线方法。结果证明,教会模型识别并应用核心技巧可显著提升其数学推理能力。
原文摘要 · Abstract (English)
Although most of the automated theorem-proving approaches depend on formal proof systems, informal theorem proving can align better with large language models' (LLMs) strength in natural language processing. In this work, we identify a primary bottleneck in informal theorem proving as a lack of insight, namely the difficulty of recognizing the core techniques required to solve complex problems. To address this, we propose $\texttt{DeepInsight}$, a unified training framework designed to cultivate this essential reasoning skill and enable LLMs to perform insightful reasoning. Our framework consists of three components: (1) $\texttt{DeepInsightTheorem}$, a hierarchical dataset that structures informal proofs by explicitly extracting core techniques and proof sketches alongside the final proof; (2) a Progressive Multi-Stage SFT strategy that mimics the human learning process, teaching the model proof writing, planning, and insight identification; and (3) $\texttt{InsightPO}$, a policy optimization method that assigns structured rewards over this insight hierarchy. Our experiments on challenging mathematical benchmarks demonstrate that this insight-aware generation strategy significantly outperforms baselines. These results demonstrate that teaching models to identify and apply core techniques can substantially improve their mathematical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。