用思维树动态验证计算,提升大模型数学解题准确率
MDToC: Metacognitive Dynamic Tree of Concepts for Boosting Mathematical Problem-Solving of Large Language Models
- 构建概念树,逐层验证每个知识点的计算
- 在多个基准上超越现有方法5%以上,最高提升7.6%
- 适合需要高精度数学推理的应用场景
尽管大语言模型在数学推理方面取得进展,但在使用传统提示技术时仍难以进行计算验证。本文提出MDToC(元认知动态概念树),采用三阶段方法:构建概念树、为每个概念生成经过验证的计算结果,并通过多数投票评估不同解法。在CHAMP、MATH和Game-of-24基准上的评估显示,GPT-4-Turbo在这些任务上分别达到58.1%、86.6%和85%的准确率,优于GoT方法5%、5.4%和4%,且无需人工设计提示。MDToC在所有基础模型上均超越现有提示方法,相比ToT提升最高达7.6%,相比GoT提升6.2%,证实了元认知计算验证是提升数学推理能力的有前景方向。
原文摘要 · Abstract (English)
Despite advances in mathematical reasoning capabilities, Large Language Models (LLMs) still struggle with calculation verification when using established prompting techniques. We present MDToC (Metacognitive Dynamic Tree of Concepts), a three-phase approach that constructs a concept tree, develops accuracy-verified calculations for each concept, and employs majority voting to evaluate competing solutions. Evaluations across CHAMP, MATH, and Game-of-24 benchmarks demonstrate our MDToC's effectiveness, with GPT-4-Turbo achieving 58.1\% on CHAMP, 86.6\% on MATH, and 85\% on Game-of-24 - outperforming GoT by 5\%, 5.4\%, and 4\% on all these tasks, respectively, without hand-engineered hints. MDToC consistently surpasses existing prompting methods across all backbone models, yielding improvements of up to 7.6\% over ToT and 6.2\% over GoT, establishing metacognitive calculation verification as a promising direction for enhanced mathematical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。