用视觉动作策略+任务规划,让机械手更聪明地完成复杂双手操作。
SViP: Sequencing Bimanual Visuomotor Policies with Object-Centric Motion Primitives
- 基于场景图拆分人类示范,生成可切换的动作基元。
- 仅需20次真实演示,就能在分布外条件下稳定执行任务。
- 适合需要强泛化能力的复杂双手操作场景。
模仿学习(IL)在复杂双手操作任务中表现优异,尤其当使用高维视觉输入时。然而,视觉动作策略的泛化能力受限,尤其是在小规模示范数据下。累积误差严重影响长时任务的完成。为此,我们提出SViP框架,将视觉动作策略无缝整合进任务与运动规划(TAMP)。SViP通过语义场景图监控将人类示范分解为双手与单手操作,利用关键场景图中的连续决策变量训练切换条件生成器。该生成器输出参数化的脚本化基元,在遭遇分布外观测时仍能保持可靠性能。仅用20次真实世界示范,SViP即可实现对分布外初始状态的泛化,无需物体位姿估计器。对于未见任务,SViP可自动发现有效解决方案,利用TAMP中的约束建模机制。真实实验表明,其性能优于现有最先进的生成式模仿学习方法,展现出更广泛的应用潜力。
原文摘要 · Abstract (English)
Imitation learning (IL), particularly when leveraging high-dimensional visual inputs for policy training, has proven intuitive and effective in complex bimanual manipulation tasks. Nonetheless, the generalization capability of visuomotor policies remains limited, especially when small demonstration datasets are available. Accumulated errors in visuomotor policies significantly hinder their ability to complete long-horizon tasks. To address these limitations, we propose SViP, a framework that seamlessly integrates visuomotor policies into task and motion planning (TAMP). SViP partitions human demonstrations into bimanual and unimanual operations using a semantic scene graph monitor. Continuous decision variables from the key scene graph are employed to train a switching condition generator. This generator produces parameterized scripted primitives that ensure reliable performance even when encountering out-of-the-distribution observations. Using only 20 real-world demonstrations, we show that SViP enables visuomotor policies to generalize across out-of-distribution initial conditions without requiring object pose estimators. For previously unseen tasks, SViP automatically discovers effective solutions to achieve the goal, leveraging constraint modeling in TAMP formulism. In real-world experiments, SViP outperforms state-of-the-art generative IL methods, indicating wider applicability for more complex tasks. Project website: https://sites.google.com/view/svip-bimanual
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。