arXiv:2411.01801cs.CVcs.LG2024-11NeurIPS被引 8

用自调节机制提升物体感知的视觉表示质量

Bootstrapping Top-down Information for Self-modulating Slot Attention

  • 引入自上而下路径,基于模型输出动态优化特征关注
  • 在多个合成与真实数据集上达到当前最优性能
  • 适合研究物体中心学习与视觉表征的学者参考

物体中心学习(OCL)旨在无监督情况下学习视觉场景中单个物体的表征,从而实现高效准确的视觉推理。传统OCL方法主要采用自下而上的方式,将同质视觉特征聚合以表示物体。然而,在复杂视觉环境中,由于物体内部视觉特征具有异质性,此类方法常表现不佳。为此,本文提出一种新颖的OCL框架,引入自上而下的路径:首先对单个物体语义进行自举(bootstrapping),再据此调制模型,使其优先关注与该语义相关的特征。通过根据自身输出动态调制模型,该自上而下路径显著提升了物体的表征质量。所提框架在多个合成与真实世界物体发现基准上均取得当前最优性能。

原文摘要 · Abstract (English)

Object-centric learning (OCL) aims to learn representations of individual objects within visual scenes without manual supervision, facilitating efficient and effective visual reasoning. Traditional OCL methods primarily employ bottom-up approaches that aggregate homogeneous visual features to represent objects. However, in complex visual environments, these methods often fall short due to the heterogeneous nature of visual features within an object. To address this, we propose a novel OCL framework incorporating a top-down pathway. This pathway first bootstraps the semantics of individual objects and then modulates the model to prioritize features relevant to these semantics. By dynamically modulating the model based on its own output, our top-down pathway enhances the representational quality of objects. Our framework achieves state-of-the-art performance across multiple synthetic and real-world object-discovery benchmarks.

物体中心学习自上而下表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。