用物体为中心的视觉表征提升机器人抓取的泛化能力
Object-Centric Representations Improve Policy Generalization in Robot Manipulation
- 将图像分割为独立物体,引入符合操作任务的先验知识
- 在光照、纹理变化下仍保持更高泛化性能,无需任务预训练
- 适合复杂真实场景中需要鲁棒视觉理解的机器人系统
视觉表征是机器人抓取策略学习与泛化的核心。现有方法依赖全局或密集特征,常将任务相关与无关信息纠缠,导致在分布外条件下表现受限。本文研究物体为中心的表征(OCR),通过将视觉输入分割为一组独立实体,引入更契合操作任务的归纳偏置。我们在一系列模拟与真实世界的抓取任务中,对比了多种视觉编码器——包括物体中心、全局与密集方法——在光照、纹理变化及干扰物存在等多样视觉条件下的泛化表现。结果表明,即使不进行任务特定预训练,基于OCR的策略在泛化任务中显著优于密集与全局表征。这表明,OCR是构建动态真实世界中高效泛化视觉系统的有前景方向。
原文摘要 · Abstract (English)
Visual representations are central to the learning and generalization capabilities of robotic manipulation policies. While existing methods rely on global or dense features, such representations often entangle task-relevant and irrelevant scene information, limiting robustness under distribution shifts. In this work, we investigate object-centric representations (OCR) as a structured alternative that segments visual input into a finished set of entities, introducing inductive biases that align more naturally with manipulation tasks. We benchmark a range of visual encoders-object-centric, global and dense methods-across a suite of simulated and real-world manipulation tasks ranging from simple to complex, and evaluate their generalization under diverse visual conditions including changes in lighting, texture, and the presence of distractors. Our findings reveal that OCR-based policies outperform dense and global representations in generalization settings, even without task-specific pretraining. These insights suggest that OCR is a promising direction for designing visual systems that generalize effectively in dynamic, real-world robotic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。