arXiv:2511.22445cs.RO2025-11被引 2

通过融合视觉与几何信息,提升机器人在复杂环境下的泛化能力。

DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization

  • 训练时用模态丢弃机制让视觉和几何分支各自保持信息量
  • 在18个仿真任务中平均性能优于基线39.1%,未知干扰下提升41.5%
  • 可零样本迁移至未见物体,实现亚厘米级空间泛化

模仿学习已成为从示范中获取视觉运动技能的关键方法,其中设计有效的观察编码器对策略泛化至关重要。然而,现有方法在测试条件与示范不一致时表现不佳,如光照、纹理、视角、物体位置或身份变化。为此,我们提出DIPOLE(基于互补编码器的扩散策略),通过训练阶段的机制融合互补模态,而非专用融合架构。模态级丢弃在每步训练中随机屏蔽一个分支,促使各模态独立保持信息性;轻量级交叉注意力层则在两者间交换互补线索。该设计使DIPOLE具备五大优势:跨多种任务稳定高性能,对视觉变化鲁棒,实现亚厘米级空间泛化,涌现出超越单一模态的能力,以及对未见物体实现零样本迁移。在18个仿真任务和4个真实世界任务中,其平均性能比六种基线高出39.1%,在未知视觉干扰下提升41.5%,在物体随机放置条件下提升15.2%。

原文摘要 · Abstract (English)

Imitation learning has emerged as a crucial approach for acquiring visuomotor skills from demonstrations, where designing effective observation encoders is essential for policy generalization. However, existing methods tend to struggle once test-time conditions differ from the demonstrations, such as changes in lighting, texture, viewpoint, object placement, or object identity. To address this challenge, we propose DIffusion POlicy with compLementarity Encoders (DIPOLE), a visuomotor policy that learns to fuse complementary modalities through a training-time mechanism rather than a specialized fusion architecture. A modality-wise dropout masks one branch at each training step, encouraging each modality to remain individually informative. A lightweight cross-attention layer then exchanges complementary cues between the two. This design endows DIPOLE with five core strengths: stable high performance across diverse tasks, robustness to visual changes, spatial generalization at sub-centimeter precision, emergent capability beyond either modality, and zero-shot transfer to unseen objects. Across 18 simulated and 4 real-world tasks, DIPOLE outperforms six baselines by 39.1% on average, with gains of 41.5% under unseen visual distractors and 15.2% under randomized object placement.

机器人模仿学习视觉运动泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。