用视觉提示统一生成推物动作,提升机器人操作灵活性。
Visual Prompt Guided Unified Pushing Policy
- 引入轻量级视觉提示机制,指导流匹配策略生成反应式推物动作。
- 在桌面上清洁任务中成功率超基线,支持多种规划场景复用。
- 适合需要灵活物体操作的机器人系统,尤其适用于视觉大模型引导的规划。
作为最简单的非抓取操作技能之一,推动物体被广泛研究为重排物体的有效手段。然而,现有方法通常依赖多步骤的预定义推物原语组合,应用范围有限,限制了其在不同场景下的效率与泛化能力。本文提出一种统一的推物策略,将轻量级提示机制融入流匹配策略中,以指导生成反应式、多模态的推物动作。视觉提示可由高层规划器指定,使推物策略能在多种规划问题中复用。实验表明,所提方法不仅优于现有基线,还能作为视觉语言模型引导规划框架中的底层原语,高效完成桌面清洁任务。
原文摘要 · Abstract (English)
As one of the simplest non-prehensile manipulation skills, pushing has been widely studied as an effective means to rearrange objects. Existing approaches, however, typically rely on multi-step push plans composed of pre-defined pushing primitives with limited application scopes, which restrict their efficiency and versatility across different scenarios. In this work, we propose a unified pushing policy that incorporates a lightweight prompting mechanism into a flow matching policy to guide the generation of reactive, multimodal pushing actions. The visual prompt can be specified by a high-level planner, enabling the reuse of the pushing policy across a wide range of planning problems. Experimental results demonstrate that the proposed unified pushing policy not only outperforms existing baselines but also effectively serves as a low-level primitive within a VLM-guided planning framework to solve table-cleaning tasks efficiently.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。