用多个互补特征原型提升无监督图像表征学习效果
Self-Organizing Visual Prototypes for Non-Parametric Representation Learning
- 用多个支持嵌入共同表征一个原型,捕捉更全面的视觉特征
- 在检索任务中达到当前最优,且模型越复杂效果越显著
- 适合研究无监督学习与非参数化表征的学者
我们提出自组织视觉原型(SOP),一种新的无监督视觉特征学习训练方法。与现有基于单个原型的自监督学习方法不同,SOP将原型由多个语义相似的支持嵌入(SEs)构成,每个支持嵌入包含互补特征,共同更准确刻画空间区域并提升训练性能。我们通过引入两种新非参数化损失函数实现该策略,并提出SOP掩码图像建模(SOP-MIM)任务,从多个非参数局部支持嵌入视角重建被遮蔽表示。我们在检索、线性评估、微调和目标检测等多类基准上全面评估了SOP学习的表征,预训练编码器在多个检索基准上达到最先进水平,且随着模型复杂度增加表现持续提升。
原文摘要 · Abstract (English)
We present Self-Organizing Visual Prototypes (SOP), a new training technique for unsupervised visual feature learning. Unlike existing prototypical self-supervised learning (SSL) methods that rely on a single prototype to encode all relevant features of a hidden cluster in the data, we propose the SOP strategy. In this strategy, a prototype is represented by many semantically similar representations, or support embeddings (SEs), each containing a complementary set of features that together better characterize their region in space and maximize training performance. We reaffirm the feasibility of non-parametric SSL by introducing novel non-parametric adaptations of two loss functions that implement the SOP strategy. Notably, we introduce the SOP Masked Image Modeling (SOP-MIM) task, where masked representations are reconstructed from the perspective of multiple non-parametric local SEs. We comprehensively evaluate the representations learned using the SOP strategy on a range of benchmarks, including retrieval, linear evaluation, fine-tuning, and object detection. Our pre-trained encoders achieve state-of-the-art performance on many retrieval benchmarks and demonstrate increasing performance gains with more complex encoders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。