通过递归分解思维,提升大模型推理能力
Recursive Decomposition of Logical Thoughts: Framework for Superior Reasoning and Knowledge Propagation in Large Language Models
- 将复杂问题逐步拆解为更小的子任务,分层推进
- 在GSM8K上达90.98%准确率,比现有方法高6.28%
- 适合需要深度逻辑推理的科研与教育场景
提升大语言模型的推理能力仍是人工智能的关键挑战。本文提出RDoLT(递归逻辑思维分解提示框架),通过三项创新显著增强模型推理性能:(1) 将复杂推理任务递归分解为逐步复杂的子任务;(2) 采用先进选择与评分机制识别最具潜力的推理路径;(3) 引入知识传播模块,模拟人类学习,追踪强弱思维以实现信息传递。在GSM8K、SVAMP、MultiArith、LastLetterConcatenation和Gaokao2023 Math等多个基准上评估,RDoLT在使用ChatGPT-4时于GSM8K达到90.98%准确率,超越现有最佳方法6.28%;其他任务中准确率提升5.5%至6.75%。结果表明RDoLT在提示工程方面具有显著潜力,为复杂推理任务提供了更高效、通用的新范式。
原文摘要 · Abstract (English)
Enhancing the reasoning capabilities of Large Language Models remains a critical challenge in artificial intelligence. We introduce RDoLT, Recursive Decomposition of Logical Thought prompting, a novel framework that significantly boosts LLM reasoning performance. RDoLT is built on three key innovations: (1) recursively breaking down complex reasoning tasks into sub-tasks of progressive complexity; (2) employing an advanced selection and scoring mechanism to identify the most promising reasoning thoughts; and (3) integrating a knowledge propagation module that mimics human learning by keeping track of strong and weak thoughts for information propagation. Our approach was evaluated across multiple benchmarks, including GSM8K, SVAMP, MultiArith, LastLetterConcatenation, and Gaokao2023 Math. The results demonstrate that RDoLT consistently outperforms existing state-of-the-art techniques, achieving a 90.98 percent accuracy on GSM8K with ChatGPT-4, surpassing state-of-the-art techniques by 6.28 percent. Similar improvements were observed on other benchmarks, with accuracy gains ranging from 5.5 percent to 6.75 percent. These findings highlight RDoLT's potential to advance prompt engineering, offering a more effective and generalizable approach to complex reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。