arXiv:2412.18743cs.CVcs.AI2024-12被引 3

对比发现物体中心模型在组合泛化上表现更好,但仍有局限。

Successes and Limitations of Object-centric Models at Compositional Generalisation

  • 用物体中心架构测试属性组合泛化能力
  • 模型能泛化到未见的物体属性组合,准确率显著提升
  • 适合研究可解释性与组合推理的开发者

近年来研究表明,标准解耦潜在变量模型在视觉领域无法实现稳健的组合学习。尽管设计初衷是将数据分解为独立变化因子,但其组合泛化能力极为有限。相比之下,物体中心架构展现出有前景的组合能力,但此前未被充分验证,且实验仅限于场景组合——即模型需泛化到新物体组合,而非新属性组合。本文证明,这类组合泛化能力同样适用于属性组合场景。此外,我们揭示了其能力来源,并通过精心训练进一步提升性能。最后,指出仍存在一个关键限制,提示未来研究方向。

原文摘要 · Abstract (English)

In recent years, it has been shown empirically that standard disentangled latent variable models do not support robust compositional learning in the visual domain. Indeed, in spite of being designed with the goal of factorising datasets into their constituent factors of variations, disentangled models show extremely limited compositional generalisation capabilities. On the other hand, object-centric architectures have shown promising compositional skills, albeit these have 1) not been extensively tested and 2) experiments have been limited to scene composition -- where models must generalise to novel combinations of objects in a visual scene instead of novel combinations of object properties. In this work, we show that these compositional generalisation skills extend to this later setting. Furthermore, we present evidence pointing to the source of these skills and how they can be improved through careful training. Finally, we point to one important limitation that still exists which suggests new directions of research.

物体中心组合泛化生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。