用扩散模型生成中间目标图,让机器人视觉伺服更灵活
Imagine2Servo: Intelligent Visual Servoing with Diffusion-Driven Goal Generation for Robotic Tasks
- 用扩散模型自动生成任务所需的中间目标图像
- 实现低重叠初始与目标图像间的长距离操控
- 支持多相机反馈,适合复杂现实场景的机器人任务
视觉伺服通过视觉传感器反馈控制机器人运动,虽已借助光流方法取得进展,但仍受限于测试时需预设目标图像、初始与目标图像重叠度要求高、依赖单摄像头反馈等问题。本文提出Imagine2Servo,利用基于扩散的图像编辑技术生成中间目标图像,突破传统限制,使视觉伺服可应用于长距离导航与操作,无需预先定义目标图像。该方法构建了以任务为导向的子目标图像合成流程,支持初始与目标图像重叠度极低的情况,并融合多摄像头反馈以完成复杂任务。实验验证了该框架在多种真实任务中的有效性与通用性,为视觉伺服系统拓展了全新能力。
原文摘要 · Abstract (English)
Visual servoing, the method of controlling robot motion through feedback from visual sensors, has seen significant advancements with the integration of optical flow-based methods. However, its application remains limited by inherent challenges, such as the necessity for a target image at test time, the requirement of substantial overlap between initial and target images, and the reliance on feedback from a single camera. This paper introduces Imagine2Servo, an innovative approach leveraging diffusion-based image editing techniques to enhance visual servoing algorithms by generating intermediate goal images. This methodology allows for the extension of visual servoing applications beyond traditional constraints, enabling tasks like long-range navigation and manipulation without predefined goal images. We propose a pipeline that synthesizes subgoal images grounded in the task at hand, facilitating servoing in scenarios with minimal initial and target image overlap and integrating multi-camera feedback for comprehensive task execution. Our contributions demonstrate a novel application of image generation to robotic control, significantly broadening the capabilities of visual servoing systems. Real-world experiments validate the effectiveness and versatility of the Imagine2Servo framework in accomplishing a variety of tasks, marking a notable advancement in the field of visual servoing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。