让大模型学会像人一样用反例检验数学定理边界,提升深层理解。
INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning

- 先内化反例构造能力,再分阶段优化策略与正确性
- 在多个模型上实现超越更大模型的数学推理表现
- 适合想提升模型深度理解力的研究者与教育应用
数学推理在大语言模型中发展迅速,但现有方法主要优化最终答案正确率,难以判断模型是否真正内化数学概念,还是仅记忆解题模式。人类数学教育中,通过构造反例来测试定理边界是深层理解的体现,但当前大模型在此方面能力薄弱。通过偏好优化提升该能力面临两大挑战:一是模型自身例子生成能力有限,难以构建有效偏好对;二是能力获取具有渐进性,需先学会采用该策略,再学会正确使用。为此,我们提出INSPIRE,一种“内化-改进”方法,结合参考引导的学生内化(RGSI),在策略模型自身分布下生成高质量偏好候选,以及分阶段评分标准偏好训练策略,将学习分解为方法导向和正确性导向两个阶段。在多个模型规模与家族上的实验表明,该方法持续提升性能,甚至超过更大的开源模型。在分布外基准测试中也验证了其未损害通用数学推理能力。
原文摘要 · Abstract (English)
Mathematical reasoning has seen rapid progress in large language models (LLMs), yet existing methods optimize predominantly for final-answer correctness, raising the question whether models truly internalize mathematical concepts or merely memorize solution patterns. In human mathematics education, example-based reasoning such as constructing counterexamples to test theorem boundaries reflects deep conceptual understanding, but remains underdeveloped in current LLMs. Enhancing this capability through preference optimization presents two key challenges: (1) the model's limited example-based reasoning ability makes constructing effective preference pairs inherently difficult; and (2) capability acquisition is progressive, as the model must first learn to adopt this strategy before learning to apply it correctly. Therefore we propose INSPIRE, an Internalize-Then-Improve approach combining Reference-Guided Student Internalization (RGSI), which produces high-quality preference candidates under the policy model's own distribution, with a stage-wise rubric preference training strategy that decomposes learning into method-oriented and correctness-oriented stages. Experiments across multiple model scales and families demonstrate consistent improvements, even surpassing larger open-source models, while evaluations on out-of-distribution benchmarks confirm no degradation in general mathematical reasoning ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。