用符号引导从无标签视频中提取可解释的视觉规划路径
Extracting Visual Plans from Unlabeled Videos via Symbolic Guidance
- 基于视觉基础模型自动提取任务符号,构建符号化状态图
- 真实机器人测试中成功率提升53%,生成速度加快35倍
- 规划过程完全可追溯,适合需要可解释性的机器人应用
视觉规划通过为条件式低层策略提供一系列中间视觉子目标,在长时程操作任务中表现优异。现有方法多依赖视频生成模型,但存在模型幻觉和计算开销大问题。本文提出Vis2Plan,一种高效、可解释且白盒的视觉规划框架,利用符号引导实现。从原始无标签的玩耍数据中,Vis2Plan借助视觉基础模型自动提取一组紧凑的任务符号,构建用于多目标、多阶段规划的高层符号转移图。测试时,给定目标任务目标,规划器在符号层面进行规划,并生成由底层符号表示支撑的物理一致的中间子目标图像序列。Vis2Plan在真实机器人环境中相比强大多扩散视频生成式视觉规划器,综合成功率提升53%,生成视觉计划速度快35倍。结果表明,Vis2Plan能生成物理一致的图像目标,同时提供完全可检查的推理步骤。
原文摘要 · Abstract (English)
Visual planning, by offering a sequence of intermediate visual subgoals to a goal-conditioned low-level policy, achieves promising performance on long-horizon manipulation tasks. To obtain the subgoals, existing methods typically resort to video generation models but suffer from model hallucination and computational cost. We present Vis2Plan, an efficient, explainable and white-box visual planning framework powered by symbolic guidance. From raw, unlabeled play data, Vis2Plan harnesses vision foundation models to automatically extract a compact set of task symbols, which allows building a high-level symbolic transition graph for multi-goal, multi-stage planning. At test time, given a desired task goal, our planner conducts planning at the symbolic level and assembles a sequence of physically consistent intermediate sub-goal images grounded by the underlying symbolic representation. Our Vis2Plan outperforms strong diffusion video generation-based visual planners by delivering 53\% higher aggregate success rate in real robot settings while generating visual plans 35$\times$ faster. The results indicate that Vis2Plan is able to generate physically consistent image goals while offering fully inspectable reasoning steps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。