用扩散模型生成相机运动,让机器人视觉伺服更稳定可靠。
DiffusionVS: A Generative Framework for Robust Visual Servoing Based on Diffusion Policy

- 以标记角点坐标为输入,通过条件去噪生成连续相机速度
- 仿真成功率近100%,实物实验达93%
- 可通用提升现有视觉伺服系统性能,适合做机器人控制
视觉伺服是机器人操作与导航的基础技术。基于回归的方法常因单步映射对噪声敏感及分布偏移导致误差累积而出现轨迹抖动。相比之下,扩散策略通过预测动作序列保持时间一致性,并借助隐式数据增强提升鲁棒性。本文提出一种新型扩散基伺服方法:基于扩散策略,以观测到的标记角点归一化图像坐标为输入,通过条件去噪生成相机速度。为克服静态数据集训练带来的泛化局限,采用在线训练范式,通过交互经验持续扩充训练数据多样性,显著提升模型性能与泛化能力。大量仿真与真实实验表明,该方法在仿真中成功率接近100%,实物实验达93%。此外,我们验证了扩散机制的通用性:将现有视觉伺服网络集成该模块后性能均得到提升。结果表明,该策略具备广泛适用性,可增强多种视觉伺服系统,不限于本文提出的特定架构。
原文摘要 · Abstract (English)
Visual servoing is a fundamental technique in robotic manipulation and navigation. Regression-based visual servoing frequently experiences trajectory jitter as a result of noise-sensitive single-step mappings and the accumulation of errors during distribution shifts. In contrast, Diffusion Policy maintains temporal consistency by predicting action sequences and improves robustness through implicit data augmentation. This paper presents a novel diffusion-based servoing method. Based on Diffusion Policy, the proposed approach uses normalized image coordinates of observed tag corners as input and generates camera velocity through conditional denoising. To overcome the generalization limitations of models trained on static datasets, an online training paradigm is adopted, continuously expanding the diversity of training data through interactive experience collection. This strategy substantially enhances both the performance and generalization capability of the model. Comprehensive simulations and real-world experiments demonstrate the effectiveness of the proposed method, achieving success rates of nearly 100\% in simulation and 93\% in physical experiments. Beyond the specific pipeline, we further validate the generality of the diffusion mechanism. Experiments show that existing visual servoing networks consistently achieve improved performance when integrated with our diffusion-based module. These results indicate that the proposed strategy possesses broad applicability and can enhance various visual servoing systems beyond the specific architecture presented here.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。