arXiv:2608.20084cs.ROcs.AI2026-08

让机器人通过视觉证据判断是否继续执行任务,避免因信息不足出错。

Evidence-Gated Task and Motion Planning with Vision-Language Models

论文配图:Evidence-Gated Task and Motion Planning with Vision-Language Models
图 1 · 摘自论文原文
  • 用视觉语言模型生成探索性子目标获取环境证据
  • 在烹饪任务中提升任务完成率,减少无效操作
  • 适合需要长时序推理与不确定环境的机器人应用

机器人在执行基于自然语言指令的长期操纵任务时,需同时考虑语义任务结构与几何可行性。但在部分可观测环境下,目标相关物体的存在可能不确定。现有结合视觉语言模型(VLM)与任务运动规划(TAMP)的方法可能依赖VLM先验知识生成子目标,缺乏观测支持,导致执行失败或意外结果。本文提出证据获取与可行性门控(EAFG)框架:通过VLM生成的探索性子目标和TAMP执行获取视觉证据,并以可行性门控决定是否继续规划、进一步采集证据或停止。实验表明,在存在物体使用模糊性的烹饪任务中,EAFG通过规划前发现任务相关物体,显著提升菜谱完成率;对于需要缺失物体的任务,能正确触发停止决策,减少对不存在物体的重复尝试。

原文摘要 · Abstract (English)

Robots executing long-horizon manipulation tasks from natural-language instructions must reason about both semantic task structure and geometric feasibility. However, under partial observability, the availability of goal-relevant objects may be uncertain. In such cases, approaches that combine Vision-Language Models (VLMs) with Task and Motion Planning (TAMP) may generate subgoals that rely on the VLM's prior knowledge without observational support, leading to execution failures or unintended outcomes. We propose Evidence Acquisition and Feasibility Gating (EAFG), a framework that acquires visual evidence through VLM-generated exploratory subgoals and TAMP-based execution. EAFG then applies a feasibility gate to decide whether to proceed with task planning, acquire further evidence, or halt. Our experiments show that, in cooking tasks with ambiguous object use, EAFG improves recipe completion by discovering task-relevant objects before planning. For instructions requiring an absent object, EAFG promotes appropriate halt decisions and reduces repeated attempts to manipulate that object.

机器人规划视觉语言模型任务与运动规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。