arXiv:2505.12500cs.AI2025-05被引 1

通过引导式探索提升大模型数学推理能力

MARGE: Improving Math Reasoning for LLMs with Guided Exploration

  • 基于自生成解题过程中的中间状态进行有目标的探索
  • 在多个基准上显著提升单次推理准确率和探索多样性
  • 无需额外标注或训练价值模型,适合希望增强推理能力的研究者

大语言模型在数学推理方面具有潜力,但其表现常受限于高质量查询数据不足。为突破此瓶颈,需通过自生成数据扩大计算量,但现有方法因各推理阶段探索无效,导致虚假相关数据泛滥。为此,我们提出MARGE:一种基于引导探索的数学推理增强方法。MARGE系统性地探索自生成解题过程中的中间状态,实现充分探索与更优的信用分配。在多种骨干模型和基准上的实验表明,MARGE无需外部标注或训练额外价值模型,即可显著提升推理能力。尤其值得注意的是,它同时提升了单次推理准确率与探索多样性,缓解了对齐方法中常见的权衡问题。结果验证了MARGE在增强数学推理及释放自生成数据潜力方面的有效性。代码与模型已开源。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit strong potential in mathematical reasoning, yet their effectiveness is often limited by a shortage of high-quality queries. This limitation necessitates scaling up computational responses through self-generated data, yet current methods struggle due to spurious correlated data caused by ineffective exploration across all reasoning stages. To address such challenge, we introduce \textbf{MARGE}: Improving \textbf{Ma}th \textbf{R}easoning with \textbf{G}uided \textbf{E}xploration, a novel method to address this issue and enhance mathematical reasoning through hit-guided exploration. MARGE systematically explores intermediate reasoning states derived from self-generated solutions, enabling adequate exploration and improved credit assignment throughout the reasoning process. Through extensive experiments across multiple backbone models and benchmarks, we demonstrate that MARGE significantly improves reasoning capabilities without requiring external annotations or training additional value models. Notably, MARGE improves both single-shot accuracy and exploration diversity, mitigating a common trade-off in alignment methods. These results demonstrate MARGE's effectiveness in enhancing mathematical reasoning capabilities and unlocking the potential of scaling self-generated training data. Our code and models are available at \href{https://github.com/georgao35/MARGE}{this link}.

数学推理大模型引导探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。