用扩散模型去除视觉干扰,让机器人模仿更稳定。
Stem-OB: Generalizable Visual Imitation Learning with Stem-Like Convergent Observation through Diffusion Inversion
- 通过扩散反演生成共性表征,消除光照纹理差异
- 真实场景成功率提升22.2%,优于最优基线
- 无需额外训练,即插即用,适合实际部署
视觉模仿学习方法表现强劲,但在面对光照、纹理等视觉扰动时泛化能力不足,限制了其在现实世界的应用。我们提出Stem-OB,利用预训练的图像扩散模型抑制低层视觉差异,同时保留高层场景结构。该图像反演过程类似于将观察结果转化为一种共享表征,其他观察均由此衍生,且去除了冗余细节。与数据增强方法不同,Stem-OB对各种未指定的外观变化具有鲁棒性,无需额外训练。本方法简单高效,可作为即插即用的解决方案。实证结果表明,该方法在仿真任务中有效,并在真实场景应用中取得显著提升,相较于最优基线平均成功率提高22.2%。
原文摘要 · Abstract (English)
Visual imitation learning methods demonstrate strong performance, yet they lack generalization when faced with visual input perturbations, including variations in lighting and textures, impeding their real-world application. We propose Stem-OB that utilizes pretrained image diffusion models to suppress low-level visual differences while maintaining high-level scene structures. This image inversion process is akin to transforming the observation into a shared representation, from which other observations stem, with extraneous details removed. Stem-OB contrasts with data-augmentation approaches as it is robust to various unspecified appearance changes without the need for additional training. Our method is a simple yet highly effective plug-and-play solution. Empirical results confirm the effectiveness of our approach in simulated tasks and show an exceptionally significant improvement in real-world applications, with an average increase of 22.2% in success rates compared to the best baseline. See https://hukz18.github.io/Stem-Ob/ for more info.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。