将开放回路动作与视觉-运动扩散策略结合,提升机器人操作的精度与速度。
Hybrid-Diffusion Models: Combining Open-loop Routines with Visuomotor Diffusion Policies
- 用预设动作原语(TAPs)增强演示数据,实现人机协同控制。
- 在真实场景中完成吸管取液、倒液和开盖等复杂任务,表现优于传统方法。
- 适合需要高精度与灵活性的机器人操作场景,如工业装配或医疗辅助。
尽管基于模仿学习的视觉-运动策略在复杂操作任务中表现良好,但通常难以达到传统控制方法的精度与速度。本文提出混合扩散模型(Hybrid-Diffusion),将开放回路规程与视觉-运动扩散策略相结合。我们设计了遥操作增强原语(TAPs),使操作者可无缝执行预定义动作,如锁定特定轴、移动至停靠位点或触发任务专属规程。该方法在推理阶段学会自动触发这些TAPs。我们在三个挑战性的真实任务中验证:试管吸取、开放式容器液体转移和容器开盖。所有实验视频均发布于项目网站:https://hybriddiffusion.github.io/
原文摘要 · Abstract (English)
Despite the fact that visuomotor-based policies obtained via imitation learning demonstrate good performances in complex manipulation tasks, they usually struggle to achieve the same accuracy and speed as traditional control based methods. In this work, we introduce Hybrid-Diffusion models that combine open-loop routines with visuomotor diffusion policies. We develop Teleoperation Augmentation Primitives (TAPs) that allow the operator to perform predefined routines, such as locking specific axes, moving to perching waypoints, or triggering task-specific routines seamlessly during demonstrations. Our Hybrid-Diffusion method learns to trigger such TAPs during inference. We validate the method on challenging real-world tasks: Vial Aspiration, Open-Container Liquid Transfer, and container unscrewing. All experimental videos are available on the project's website: https://hybriddiffusion.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。