arXiv:2512.24103cs.LGcs.AI2025-12被引 8

让大模型自我批评,显著提升规划任务表现

Enhancing LLM Planning Capabilities through Intrinsic Self-Critique

  • 用自洽式自我批判机制,无需外部验证
  • 在Blocksworld等3个数据集上超越强基线
  • 适用于各类大模型,尤其适合复杂规划场景

我们提出一种让大语言模型(LLM)自我批判其自身答案的方法,旨在提升其在规划任务中的表现。尽管先前研究质疑了自我批判的有效性,但本工作在Blocksworld、Logistics和Mini-grid三个规划数据集上,通过内在自我批判机制实现了显著性能提升,且不依赖外部验证器。我们采用少样本学习作为基础方法,并逐步扩展为多样本策略,再引入迭代修正与优化过程,进一步大幅改进性能。实验结果表明,该方法在2024年10月发布的各类大模型检查点中达到了新的最佳水平。研究重点在于方法本身的可迁移性,证明了模型具备内在自我优化能力,未来应用于更复杂的搜索策略和更强模型时有望取得更好效果。

原文摘要 · Abstract (English)

We demonstrate an approach for LLMs to critique their \emph{own} answers with the goal of enhancing their performance that leads to significant improvements over established planning benchmarks. Despite the findings of earlier research that has cast doubt on the effectiveness of LLMs leveraging self critique methods, we show significant performance gains on planning datasets in the Blocksworld domain through intrinsic self-critique, without external source such as a verifier. We also demonstrate similar improvements on Logistics and Mini-grid datasets, exceeding strong baseline accuracies. We employ a few-shot learning technique and progressively extend it to a many-shot approach as our base method and demonstrate that it is possible to gain substantial improvement on top of this already competitive approach by employing an iterative process for correction and refinement. We illustrate how self-critique can significantly boost planning performance. Our empirical results present new state-of-the-art on the class of models considered, namely LLM model checkpoints from October 2024. Our primary focus lies on the method itself, demonstrating intrinsic self-improvement capabilities that are applicable regardless of the specific model version, and we believe that applying our method to more complex search techniques and more capable models will lead to even better performance.

大模型自我批判规划推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。