arXiv:2505.11719cs.ROcs.AI2025-05被引 5

让机器人在未见视觉环境下零样本适应,提升真实场景泛化能力

Zero-Shot Visual Generalization in Robot Manipulation

  • 采用解耦表征与联想记忆结合的方法,增强视觉策略鲁棒性
  • 在仿真和真实硬件上实现零样本视觉扰动适应,效果优于现有方法
  • 新提出旋转等变技术,使策略对相机视角变化也具备抗性

基于视觉的机器人操作策略在多样化视觉环境中的鲁棒性仍是关键挑战。当前方法常依赖点云、深度等不变表示,或通过视觉域随机化和大规模数据强行泛化。解耦表征学习(尤其结合联想记忆)近期展现出在视觉分布偏移下保持鲁棒性的潜力,但多限于简单基准任务。本文将该方法扩展至更复杂视觉与动态的操作任务,在仿真与真实硬件上均实现对视觉扰动的零样本适应。进一步将该方法应用于扩散策略(Diffusion Policy),实证显示其在视觉泛化上显著优于当前最优模仿学习方法。最后,引入源自模型等变性的新技巧,可将任意神经网络策略转化为对二维平面旋转不变的版本,使策略兼具视觉与相机视角扰动的鲁棒性。本工作为实现开箱即用且适应真实世界复杂性的操作策略迈出重要一步。

原文摘要 · Abstract (English)

Training vision-based manipulation policies that are robust across diverse visual environments remains an important and unresolved challenge in robot learning. Current approaches often sidestep the problem by relying on invariant representations such as point clouds and depth, or by brute-forcing generalization through visual domain randomization and/or large, visually diverse datasets. Disentangled representation learning - especially when combined with principles of associative memory - has recently shown promise in enabling vision-based reinforcement learning policies to be robust to visual distribution shifts. However, these techniques have largely been constrained to simpler benchmarks and toy environments. In this work, we scale disentangled representation learning and associative memory to more visually and dynamically complex manipulation tasks and demonstrate zero-shot adaptability to visual perturbations in both simulation and on real hardware. We further extend this approach to imitation learning, specifically Diffusion Policy, and empirically show significant gains in visual generalization compared to state-of-the-art imitation learning methods. Finally, we introduce a novel technique adapted from the model equivariance literature that transforms any trained neural network policy into one invariant to 2D planar rotations, making our policy not only visually robust but also resilient to certain camera perturbations. We believe that this work marks a significant step towards manipulation policies that are not only adaptable out of the box, but also robust to the complexities and dynamical nature of real-world deployment. Supplementary videos are available at https://sites.google.com/view/vis-gen-robotics/home.

机器人操控视觉泛化扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。