arXiv:2503.11565cs.CVcs.RO2025-03被引 2

提出解耦的物体中心表征,提升机器人在多物环境中的抓取泛化能力。

Disentangled Object-Centric Image Representation for Robotic Manipulation

  • 将目标物、障碍物和机器人本体解耦表征,增强视觉学习的归纳偏置
  • 在多物场景中实现最优抓取性能,测试时可泛化至新目标与干扰物
  • 兼具仿真与真实世界零样本迁移能力,适合复杂交互任务

从视觉中学习机器人操作技能是实现广泛泛化的真实世界应用的有前景方法。尽管物体中心表征已被证明能提供更好的归纳偏置,从而提升性能与泛化能力,但我们发现其在多物体环境中仍难以学习简单操作技能。为此,我们提出DOCIR,一种引入目标物、障碍物与机器人本体解耦表征的物体中心框架。实验表明,该方法在多物体环境下从视觉输入学习抓取与放置技能方面达到当前最优表现,并能在测试时泛化至变化的目标物与干扰物。此外,该方法在仿真环境中有效,并实现了零样本迁移到真实世界。

原文摘要 · Abstract (English)

Learning robotic manipulation skills from vision is a promising approach for developing robotics applications that can generalize broadly to real-world scenarios. As such, many approaches to enable this vision have been explored with fruitful results. Particularly, object-centric representation methods have been shown to provide better inductive biases for skill learning, leading to improved performance and generalization. Nonetheless, we show that object-centric methods can struggle to learn simple manipulation skills in multi-object environments. Thus, we propose DOCIR, an object-centric framework that introduces a disentangled representation for objects of interest, obstacles, and robot embodiment. We show that this approach leads to state-of-the-art performance for learning pick and place skills from visual inputs in multi-object environments and generalizes at test time to changing objects of interest and distractors in the scene. Furthermore, we show its efficacy both in simulation and zero-shot transfer to the real world.

机器人操作物体中心解耦表征视觉学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。