改进思维链方法,让大模型自主生成长序列规划
LLMs Can Plan Only If We Tell Them
- 提出AoT+方法,增强大模型独立规划能力
- 在标准规划基准上超越人类表现
- 无需外部反馈,实现全自动规划
大型语言模型(LLMs)在自然语言处理和推理方面展现出显著能力,但在自主规划方面的有效性仍存争议。现有研究虽已利用外部反馈机制或在受控环境中进行规划,但这些方法常需大量计算与开发资源,且依赖精心设计与迭代回溯提示。即使最先进的模型如GPT-4,在无额外支持的情况下,也难以在典型规划基准(如Blocksworld)上达到人类水平。本文探究了大模型是否能独立生成长时序规划并媲美人类基准。通过对其算法思维链(Algorithm-of-Thoughts, AoT)的创新改进,我们提出的AoT+方法在规划基准上实现了当前最优结果,且完全自主运行,超越了先前方法与人类基线。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated significant capabilities in natural language processing and reasoning, yet their effectiveness in autonomous planning has been under debate. While existing studies have utilized LLMs with external feedback mechanisms or in controlled environments for planning, these approaches often involve substantial computational and development resources due to the requirement for careful design and iterative backprompting. Moreover, even the most advanced LLMs like GPT-4 struggle to match human performance on standard planning benchmarks, such as the Blocksworld, without additional support. This paper investigates whether LLMs can independently generate long-horizon plans that rival human baselines. Our novel enhancements to Algorithm-of-Thoughts (AoT), which we dub AoT+, help achieve state-of-the-art results in planning benchmarks out-competing prior methods and human baselines all autonomously.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。