用统一量化机制提升视觉基础模型在物体中心学习中的表现。
Vector-Quantized Vision Foundation Models for Object-Centric Learning
- 共享量化视觉基础模型特征,统一物体聚合与重建过程。
- 在物体发现、识别及下游任务中均超越现有基线方法。
- 适合研究物体中心学习与视觉基础模型融合的学者。
物体中心学习(OCL)将图像或视频特征图聚合为物体级别的特征向量,称为“槽”(slots)。现有自监督重建方法在复杂纹理物体上表现不佳,因此引入视觉基础模型(VFM)表示作为聚合输入和重建目标。尽管已有方法以不同方式利用VFM表示,但未能充分挖掘其潜力。为此,我们提出统一架构VQ-VFM-OCL(VVO),核心在于在OCL的聚合与解码过程中共享量化VFM表示。实验表明,无论使用何种VFM、聚合器或解码器,我们的VVO在物体发现、识别以及下游视觉预测与推理任务中均持续优于基线。我们还从数学上分析了为什么VFM表示有助于OCL聚合,以及为何共享量化作为重建目标能增强监督信号。代码与模型检查点已公开于https://github.com/Genera1Z/VQ-VFM-OCL。
原文摘要 · Abstract (English)
Object-Centric Learning (OCL) aggregates image or video feature maps into object-level feature vectors, termed \textit{slots}. It's self-supervision of reconstructing the input from slots struggles with complex object textures, thus Vision Foundation Model (VFM) representations are used as the aggregation input and reconstruction target. Existing methods leverage VFM representations in diverse ways yet fail to fully exploit their potential. In response, we propose a unified architecture, Vector-Quantized VFMs for OCL (VQ-VFM-OCL, or VVO). The key to our unification is simply shared quantizing VFM representations in OCL aggregation and decoding. Experiments show that across different VFMs, aggregators and decoders, our VVO consistently outperforms baselines in object discovery and recognition, as well as downstream visual prediction and reasoning. We also mathematically analyze why VFM representations facilitate OCL aggregation and why their shared quantization as reconstruction targets strengthens OCL supervision. Our source code and model checkpoints are available on https://github.com/Genera1Z/VQ-VFM-OCL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。