arXiv:2504.07092cs.CVcs.AI2025-04被引 10

用分割编码物体,比传统槽方法更擅长处理分布外泛化。

Are We Done with Object-Centric Learning?

  • 通过像素级分割独立编码物体,无需训练即可实现
  • 在分布外物体发现任务中显著优于传统槽方法
  • 适合研究物体表征与人类认知的学者使用

对象中心学习(OCL)旨在学习仅包含单一对象、与场景中其他对象或背景无关的表示。这一目标支撑了分布外(OOD)泛化、样本高效组合及结构化环境建模等任务。以往研究多聚焦于无监督机制,将对象分离至表示空间中的离散槽,评估方式为无监督对象发现。然而,随着近期高效分割模型的发展,我们可在像素空间中分离对象并独立编码,从而在零样本条件下实现优异的分布外对象发现性能,且可扩展至基础模型,天然支持可变数量的槽。因此,获得对象中心表示的目标已基本达成。但关键问题仍存:在场景中分离对象如何促进更广泛的OCL目标,如分布外泛化?为此,我们从OCL视角探究由虚假背景线索引发的泛化挑战,提出一种无需训练的探测器——基于掩码的对象中心分类(OCCAM)。实验证明,基于分割的编码显著优于槽式OCL方法。尽管如此,实际应用仍面临挑战。我们为OCL社区提供工具箱,推动可扩展表示在实践中的应用,并关注人类认知中对象感知等根本问题。代码已公开:https://github.com/AlexanderRubinstein/OCCAM。

原文摘要 · Abstract (English)

Object-centric learning (OCL) seeks to learn representations that only encode an object, isolated from other objects or background cues in a scene. This approach underpins various aims, including out-of-distribution (OOD) generalization, sample-efficient composition, and modeling of structured environments. Most research has focused on developing unsupervised mechanisms that separate objects into discrete slots in the representation space, evaluated using unsupervised object discovery. However, with recent sample-efficient segmentation models, we can separate objects in the pixel space and encode them independently. This achieves remarkable zero-shot performance on OOD object discovery benchmarks, is scalable to foundation models, and can handle a variable number of slots out-of-the-box. Hence, the goal of OCL methods to obtain object-centric representations has been largely achieved. Despite this progress, a key question remains: How does the ability to separate objects within a scene contribute to broader OCL objectives, such as OOD generalization? We address this by investigating the OOD generalization challenge caused by spurious background cues through the lens of OCL. We propose a novel, training-free probe called Object-Centric Classification with Applied Masks (OCCAM), demonstrating that segmentation-based encoding of individual objects significantly outperforms slot-based OCL methods. However, challenges in real-world applications remain. We provide the toolbox for the OCL community to use scalable object-centric representations, and focus on practical applications and fundamental questions, such as understanding object perception in human cognition. Our code is available here: https://github.com/AlexanderRubinstein/OCCAM.

对象中心分割编码分布外泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。