提升物体中心学习的语义感知与灵活性,让模型更懂物体形状和关系。
ContextFusion and Bootstrap: An Effective Approach to Improve Slot Attention-Based Object-Centric Learning
- 引入上下文融合模块,结合前景背景语义信息增强特征
- 设计自举分支实现特征自适应,支持编码器灵活微调
- 在模拟与真实数据集上显著提升主流槽注意力模型表现
人类能将场景分解为独立物体,并通过其关系理解环境。物体中心学习旨在无监督地模拟这一过程。近年来,基于槽注意力的框架成为主流,广泛应用于各类下游任务。然而现有方法存在两大局限:(1) 缺乏高层语义信息,当前方法仅依赖颜色、纹理等低层特征分配图像区域到槽,导致对物体轮廓、形状等语义理解不足;(2) 无法微调编码器,因需保持训练全程特征空间稳定以实现从槽重建,限制了有效学习的灵活性。为此,我们提出一种新的上下文融合阶段和自举分支,可无缝集成至现有槽注意力模型。在上下文融合阶段,利用前景与背景的语义信息,并引入辅助指示符提供额外上下文线索,丰富超出低层特征的语义内容。在自举分支中,将特征适应与原始重建阶段解耦,采用自举策略训练特征自适应机制,实现更高灵活性。实验表明,该方法显著提升了多种SOTA槽注意力模型在模拟与真实数据集上的性能。
原文摘要 · Abstract (English)
A key human ability is to decompose a scene into distinct objects and use their relationships to understand the environment. Object-centric learning aims to mimic this process in an unsupervised manner. Recently, the slot attention-based framework has emerged as a leading approach in this area and has been widely used in various downstream tasks. However, existing slot attention methods face two key limitations: (1) a lack of high-level semantic information. In current methods, image areas are assigned to slots based on low-level features such as color and texture. This makes the model overly sensitive to low-level features and limits its understanding of object contours, shapes, or other semantic characteristics. (2) The inability to fine-tune the encoder. Current methods require a stable feature space throughout training to enable reconstruction from slots, which restricts the flexibility needed for effective object-centric learning. To address these limitations, we propose a novel ContextFusion stage and a Bootstrap Branch, both of which can be seamlessly integrated into existing slot attention models. In the ContextFusion stage, we exploit semantic information from the foreground and background, incorporating an auxiliary indicator that provides additional contextual cues about them to enrich the semantic content beyond low-level features. In the Bootstrap Branch, we decouple feature adaptation from the original reconstruction phase and introduce a bootstrap strategy to train a feature-adaptive mechanism, allowing for more flexible adaptation. Experimental results show that our method significantly improves the performance of different SOTA slot attention models on both simulated and real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。