用反例训练大模型,提升数学推理能力
One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
- 设计反例驱动的数学推理任务,检验模型概念理解
- 构建大学级反例基准CounterMATH,挑战主流模型
- 发现当前模型缺乏反例推理能力,需专项训练
利用数学大语言模型(LLMs)生成证明是当前大模型研究的核心议题。我们指出,现有模型的证明能力高度依赖于训练中是否接触过相关证明过程,这限制了其对数学定理及概念的深层理解。受人类数学教育中“反证法”教学方法的启发,本文旨在通过反例增强大模型的数学推理与证明能力。具体而言,我们手动构建了一个高质量、大学水平的数学基准测试集CounterMATH,要求模型通过提供反例来证明数学命题,以评估其对数学概念的掌握程度。同时,我们开发了一套数据工程框架,用于自动生成训练数据以持续优化模型。大量实验与分析表明,CounterMATH具有挑战性,例如OpenAI o1等主流模型在反例驱动的证明任务上表现不足。进一步研究发现,强化模型的反例驱动概念推理能力,对提升其整体数学能力至关重要。本工作为数学大模型社区提供了新视角。
原文摘要 · Abstract (English)
Leveraging mathematical Large Language Models (LLMs) for proof generation is a fundamental topic in LLMs research. We argue that the ability of current LLMs to prove statements largely depends on whether they have encountered the relevant proof process during training. This reliance limits their deeper understanding of mathematical theorems and related concepts. Inspired by the pedagogical method of "proof by counterexamples" commonly used in human mathematics education, our work aims to enhance LLMs' ability to conduct mathematical reasoning and proof through counterexamples. Specifically, we manually create a high-quality, university-level mathematical benchmark, CounterMATH, which requires LLMs to prove mathematical statements by providing counterexamples, thereby assessing their grasp of mathematical concepts. Additionally, we develop a data engineering framework to automatically obtain training data for further model improvement. Extensive experiments and detailed analyses demonstrate that CounterMATH is challenging, indicating that LLMs, such as OpenAI o1, have insufficient counterexample-driven proof capabilities. Moreover, our exploration into model training reveals that strengthening LLMs' counterexample-driven conceptual reasoning abilities is crucial for improving their overall mathematical capabilities. We believe that our work offers new perspectives on the community of mathematical LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。