arXiv:2602.01662cs.RO2026-02被引 3

让机器人通过计划语言一步步验证执行,实现复杂任务的自主操作。

PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation

  • 用物体状态、动作模板和符号化计划构建可验证的推理空间。
  • 每步操作后通过视觉反馈检查预期效果,失败时自动重规划。
  • 适用于多种机器人和任务,揭示当前视觉语言模型的局限性。

视觉语言模型(VLMs)的进步推动了真实世界机器人操作的发展。然而,在非结构化环境中进行长时程操作仍需模型对场景状态变化、动作约束及执行结果进行推理,仅靠自然语言推理难以应对。我们提出PLanAR,一种基于计划语言的机器人智能体框架,支持开放词汇、长时程操作。该框架通过计划语言接口定义VLM的推理空间:物体谓词表示场景状态,动作模式定义带前提与效果的机器人技能,符号化计划提供可执行的中间表示。这一接口支持分步验证:每次动作后,系统利用本地观测检查预期符号效果是否达成,从而更新任务状态、检测失败并重新规划。在不同机器人平台、VLM后端及多项任务(包括堆叠、填字游戏、长流程厨房任务)中,PLanAR展现出强大的现实能力,同时揭示了当前VLM在具身推理中的关键不足。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have enabled increasing progress in real-world robot manipulation. However, long-horizon manipulation in unstructured environments requires VLMs to reason about changing scene states, action constraints, and execution outcomes, which remains difficult with natural language reasoning alone. We present PLanAR, a planning-language-grounded robot agent framework for open-vocabulary, long-horizon manipulation. PLanAR uses a planning-language interface to define the VLM reasoning space: object predicates represent scene states, action schemas specify robot skills with preconditions and effects, and symbolic plans provide executable intermediate representations. This interface enables stepwise verification: after each action, PLanAR uses onboard observations to check whether the expected symbolic effects have been achieved, allowing the VLM-based agent to update task states, detect failures, and replan when execution deviates from expectation. Across robot embodiments, VLM backends, and tasks including stacking, crossword solving, and long-horizon kitchen workflows, PLanAR demonstrates strong real-world capability while revealing key limitations of current VLMs in embodied reasoning.

机器人操作计划推理视觉语言模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。