arXiv:2602.16689cs.CVcs.LG2026-02

对比了物体中心表示在组合泛化上的表现,发现其在数据少或计算资源有限时更优。

Are Object-Centric Representations Better At Compositional Generalization?

  • 用三个可控视觉世界测试物体中心与密集表示的泛化能力
  • 物体中心方法在复杂组合任务中表现更好,且需更少训练数据
  • 当数据量充足时,密集表示才可追上,但需更多下游计算

组合泛化是人类认知的核心能力,也是机器学习的关键挑战。物体中心(OC)表示将场景编码为一组物体,常被认为能支持此类泛化,但在视觉丰富场景中的系统性证据仍不足。我们引入一个跨三个受控视觉世界(CLEVRTex、Super-CLEVR、MOVi-C)的视觉问答基准,评估带与不带物体中心偏见的视觉编码器对未见过的对象属性组合的泛化能力。为确保公平比较,我们控制了训练数据多样性、样本量、表示规模、下游模型容量和计算资源。以DINOv2和SigLIP2两类广泛使用的视觉编码器及其对应的OC版本为基础模型。关键发现包括:(1) 在更复杂的组合泛化设置中,OC方法表现更优;(2) 原始密集表示仅在简单任务中胜出,且通常需要显著更多的下游计算;(3) OC模型更具样本效率,在较少图像下即实现更强泛化,而密集编码器仅在数据充分且多样时才能追上甚至超越。总体而言,当数据集规模、训练数据多样性或下游计算受限时,物体中心表示提供更强的组合泛化能力。

原文摘要 · Abstract (English)

Compositional generalization, the ability to reason about novel combinations of familiar concepts, is fundamental to human cognition and a critical challenge for machine learning. Object-centric (OC) representations, which encode a scene as a set of objects, are often argued to support such generalization, but systematic evidence in visually rich settings is limited. We introduce a Visual Question Answering benchmark across three controlled visual worlds (CLEVRTex, Super-CLEVR, and MOVi-C) to measure how well vision encoders, with and without object-centric biases, generalize to unseen combinations of object properties. To ensure a fair and comprehensive comparison, we carefully account for training data diversity, sample size, representation size, downstream model capacity, and compute. We use DINOv2 and SigLIP2, two widely used vision encoders, as the foundation models and their OC counterparts. Our key findings reveal that (1) OC approaches are superior in harder compositional generalization settings; (2) original dense representations surpass OC only on easier settings and typically require substantially more downstream compute; and (3) OC models are more sample efficient, achieving stronger generalization with fewer images, whereas dense encoders catch up or surpass them only with sufficient data and diversity. Overall, object-centric representations offer stronger compositional generalization when any one of dataset size, training data diversity, or downstream compute is constrained.

组合泛化物体中心视觉编码器样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。