arXiv:2506.19408cs.AIcs.RO2025-06被引 3

对比了物体中心表示在机器人操作任务中的表现,发现其更适应复杂场景。

Is an object-centric representation beneficial for robotic manipulation ?

  • 用模拟环境中的多物体操作任务测试物体中心表示方法
  • 在复杂场景下,物体中心方法比整体表示更少失败
  • 适合研究数据效率和泛化能力的视觉学习方向

物体中心表示(OCR)近年来在计算机视觉领域受到关注,被认为有助于学习图像和视频的结构化表征,并提升下游任务的数据效率与泛化能力。然而,现有工作大多仅在场景分解任务上评估,缺乏对所学表示的推理能力检验。机器人操作通常涉及多物体环境及潜在的物体间交互,是验证物体中心表示潜力的理想场景。为此,我们在模拟环境中构建多个包含多个物体(如干扰物、机器人等)且高度随机化的操作任务(物体位置、颜色、形状、背景、初始状态等均随机)。评估一种经典物体中心方法在多种泛化场景下的表现,并与若干最先进的整体表示方法进行对比。结果表明,在复杂场景结构下,现有方法易出现失败,而物体中心方法能有效克服这些挑战。

原文摘要 · Abstract (English)

Object-centric representation (OCR) has recently become a subject of interest in the computer vision community for learning a structured representation of images and videos. It has been several times presented as a potential way to improve data-efficiency and generalization capabilities to learn an agent on downstream tasks. However, most existing work only evaluates such models on scene decomposition, without any notion of reasoning over the learned representation. Robotic manipulation tasks generally involve multi-object environments with potential inter-object interaction. We thus argue that they are a very interesting playground to really evaluate the potential of existing object-centric work. To do so, we create several robotic manipulation tasks in simulated environments involving multiple objects (several distractors, the robot, etc.) and a high-level of randomization (object positions, colors, shapes, background, initial positions, etc.). We then evaluate one classical object-centric method across several generalization scenarios and compare its results against several state-of-the-art hollistic representations. Our results exhibit that existing methods are prone to failure in difficult scenarios involving complex scene structures, whereas object-centric methods help overcome these challenges.

物体中心机器人操作泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。