用蒙特卡洛树搜索提升大模型数学推理能力,效果超越现有方法。
MC-NEST: Enhancing Mathematical Reasoning in Large Language Models leveraging a Monte Carlo Self-Refine Tree
- 基于大模型自反思与自评估的蒙特卡洛树搜索优化策略。
- GPT-4o在AIME上达到38.6分,MathOdyssey上达12.6分,性能领先。
- 适用于需要多步逻辑推演的数学竞赛题,适合推理研究者使用。
大语言模型在数学推理方面面临重大挑战。为此,我们提出蒙特卡洛自修正树(MC-NEST),一种融合大模型自反思与自评估的蒙特卡洛树搜索扩展方法,以增强复杂推理任务中的决策能力。MC-NEST利用上限置信区间(UCT)分数结合多样选择策略,平衡探索与利用。通过迭代式批判与优化,大模型学会更战略性的推理方式。实验证明,采用重要性采样策略的MC-NEST显著提升GPT-4o性能,在奥数级基准测试中达到当前最优的pass@1分数:AIME为38.6,MathOdyssey为12.6。使用GPT-4o和Phi-3-mini时,解题质量分别达到84.0%和82.08%,表现出跨模型的强一致性。MC-NEST在代数、几何与数论任务中表现优异,得益于其处理抽象、逻辑演绎及多步推理的核心能力。
原文摘要 · Abstract (English)
Mathematical reasoning presents significant challenges for large language models (LLMs). To enhance their capabilities, we propose Monte Carlo Self-Refine Tree (MC-NEST), an extension of Monte Carlo Tree Search that integrates LLM-based self-refinement and self-evaluation for improved decision-making in complex reasoning tasks. MC-NEST balances exploration and exploitation using Upper Confidence Bound (UCT) scores combined with diverse selection policies. Through iterative critique and refinement, LLMs learn to reason more strategically. Empirical results demonstrate that MC-NEST with an importance sampling policy substantially improves GPT-4o's performance, achieving state-of-the-art pass@1 scores on Olympiad-level benchmarks. Specifically, MC-NEST attains a pass@1 of 38.6 on AIME and 12.6 on MathOdyssey. The solution quality for MC-NEST using GPT-4o and Phi-3-mini reaches 84.0\% and 82.08\%, respectively, indicating robust consistency across different LLMs. MC-NEST performs strongly across Algebra, Geometry, and Number Theory, benefiting from its ability to handle abstraction, logical deduction, and multi-step reasoning -- core skills in mathematical problem solving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。