arXiv:2509.09790cs.AI2025-09被引 2

大模型能有效提供规划反馈,减少对人工设计奖励的依赖。

How well can LLMs provide planning feedback in grounded environments?

  • 用大模型生成多种规划反馈,如动作建议、目标指导等。
  • 模型越大、推理越强,反馈越准确且偏差越小。
  • 在复杂动态或连续空间环境中反馈质量下降。

在具身环境中学习规划通常需要精心设计的奖励函数或高质量的标注示范。近期研究发现,预训练的基础模型(如大语言模型LLMs和视觉语言模型VLMs)具备有助于规划的背景知识,从而减少了政策学习所需的奖励设计和示范数量。本文评估了LLMs和VLMs在符号、语言和连续控制环境中的反馈能力,涵盖二元反馈、偏好反馈、动作建议、目标建议和动作增量反馈等多种类型。同时考察了上下文学习、思维链和环境动态访问等推理方法对反馈性能的影响。结果表明,基础模型可在多领域提供多样且高质量的反馈;更大的模型和具备推理能力的模型始终表现更优,误差更小,且更受益于增强的推理方法。然而,在具有复杂动态或连续状态与动作空间的环境中,反馈质量会下降。

原文摘要 · Abstract (English)

Learning to plan in grounded environments typically requires carefully designed reward functions or high-quality annotated demonstrations. Recent works show that pretrained foundation models, such as large language models (LLMs) and vision language models (VLMs), capture background knowledge helpful for planning, which reduces the amount of reward design and demonstrations needed for policy learning. We evaluate how well LLMs and VLMs provide feedback across symbolic, language, and continuous control environments. We consider prominent types of feedback for planning including binary feedback, preference feedback, action advising, goal advising, and delta action feedback. We also consider inference methods that impact feedback performance, including in-context learning, chain-of-thought, and access to environment dynamics. We find that foundation models can provide diverse high-quality feedback across domains. Moreover, larger and reasoning models consistently provide more accurate feedback, exhibit less bias, and benefit more from enhanced inference methods. Finally, feedback quality degrades for environments with complex dynamics or continuous state spaces and action spaces.

大模型规划反馈具身智能强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。