arXiv:2509.02761cs.AI2025-09被引 5

用大模型自动修正机器人任务计划中的冗余和错误,提升执行效率。

Plan Verification for LLM-Based Embodied Task Completion Agents

  • 通过判官与规划者大模型迭代纠错,生成更清晰的动作序列。
  • 在四个主流大模型上实现90%召回率和100%精确率,96.5%的计划三轮内完成优化。
  • 保留人类纠错模式,适合用于提升具身智能的模仿学习数据质量。

基于大语言模型(LLM)的任务计划与人类示范在具身人工智能中可能包含噪声,如多余动作、重复导航和逻辑错误,影响策略质量。本文提出一种迭代验证框架:由判官LLM批评动作序列,规划者LLM据此修正,逐步生成更清洁、空间连贯的轨迹。与规则方法不同,该方法依赖自然语言提示,可泛化处理无关动作、矛盾和缺失步骤等多种错误。在TEACh具身AI数据集的人工标注动作上,该框架在四个先进大模型(GPT o4-mini、DeepSeek-R1、Gemini 2.5、LLaMA 4 Scout)上达到最高90%召回率与100%精确率。优化循环收敛迅速,96.5%的序列最多经三轮迭代完成;同时提升时间效率与空间动作组织性。关键的是,该方法保留人类纠错行为模式,而非将其消除,为未来鲁棒纠错行为研究提供支持。本工作确立了计划验证作为大模型在空间规划与动作优化中的可靠能力,为具身智能模仿学习提供了可扩展的高质量训练数据路径。

原文摘要 · Abstract (English)

Large language model (LLM) based task plans and corresponding human demonstrations for embodied AI may be noisy, with unnecessary actions, redundant navigation, and logical errors that reduce policy quality. We propose an iterative verification framework in which a Judge LLM critiques action sequences and a Planner LLM applies the revisions, yielding progressively cleaner and more spatially coherent trajectories. Unlike rule-based approaches, our method relies on natural language prompting, enabling broad generalization across error types including irrelevant actions, contradictions, and missing steps. On a set of manually annotated actions from the TEACh embodied AI dataset, our framework achieves up to 90% recall and 100% precision across four state-of-the-art LLMs (GPT o4-mini, DeepSeek-R1, Gemini 2.5, LLaMA 4 Scout). The refinement loop converges quickly, with 96.5% of sequences requiring at most three iterations, while improving both temporal efficiency and spatial action organization. Crucially, the method preserves human error-recovery patterns rather than collapsing them, supporting future work on robust corrective behavior. By establishing plan verification as a reliable LLM capability for spatial planning and action refinement, we provide a scalable path to higher-quality training data for imitation learning in embodied AI.

具身智能大模型任务规划验证框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。