arXiv:2505.21533cs.CVcs.LG2025-05ICML被引 2

用多个互补特征原型提升无监督图像表征学习效果

Self-Organizing Visual Prototypes for Non-Parametric Representation Learning

  • 用多个支持嵌入共同表征一个原型,捕捉更全面的视觉特征
  • 在检索任务中达到当前最优,且模型越复杂效果越显著
  • 适合研究无监督学习与非参数化表征的学者

我们提出自组织视觉原型(SOP),一种新的无监督视觉特征学习训练方法。与现有基于单个原型的自监督学习方法不同,SOP将原型由多个语义相似的支持嵌入(SEs)构成,每个支持嵌入包含互补特征,共同更准确刻画空间区域并提升训练性能。我们通过引入两种新非参数化损失函数实现该策略,并提出SOP掩码图像建模(SOP-MIM)任务,从多个非参数局部支持嵌入视角重建被遮蔽表示。我们在检索、线性评估、微调和目标检测等多类基准上全面评估了SOP学习的表征,预训练编码器在多个检索基准上达到最先进水平,且随着模型复杂度增加表现持续提升。

原文摘要 · Abstract (English)

We present Self-Organizing Visual Prototypes (SOP), a new training technique for unsupervised visual feature learning. Unlike existing prototypical self-supervised learning (SSL) methods that rely on a single prototype to encode all relevant features of a hidden cluster in the data, we propose the SOP strategy. In this strategy, a prototype is represented by many semantically similar representations, or support embeddings (SEs), each containing a complementary set of features that together better characterize their region in space and maximize training performance. We reaffirm the feasibility of non-parametric SSL by introducing novel non-parametric adaptations of two loss functions that implement the SOP strategy. Notably, we introduce the SOP Masked Image Modeling (SOP-MIM) task, where masked representations are reconstructed from the perspective of multiple non-parametric local SEs. We comprehensively evaluate the representations learned using the SOP strategy on a range of benchmarks, including retrieval, linear evaluation, fine-tuning, and object detection. Our pre-trained encoders achieve state-of-the-art performance on many retrieval benchmarks and demonstrate increasing performance gains with more complex encoders.

无监督学习视觉表征非参数化原型学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。