arXiv:2501.02860cs.LGcs.CV2025-01被引 2

通过对齐局部与全局特征,提升自监督学习的性能与鲁棒性。

Seeing the Whole in the Parts in Self-Supervised Representation Learning

  • 用局部表示与全局图像表示对齐建模空间关系
  • ImageNet-1K上达71.5%准确率,优于此前方法
  • 对噪声、对抗攻击和大裁剪更鲁棒,适合工业级应用

自监督学习(SSL)近期通过遮蔽图像部分或大幅裁剪来建模视觉特征的空间共现。本文提出一种新方法——CO-SSL,通过将局部表示(池化前)与全局图像表示对齐来建模空间共现。该方法为实例判别类算法家族,在多个数据集上表现优于以往方法,尤其在ImageNet-1K上以100个预训练周期达到71.5%的Top-1准确率。此外,CO-SSL在面对噪声污染、内部干扰、小规模对抗攻击及大训练裁剪尺寸时均表现出更强鲁棒性。分析表明,其学习到的高度冗余局部表示可能是鲁棒性的来源。整体上,本研究提示:对齐局部与全局表示或可成为无监督类别学习的关键原则。

原文摘要 · Abstract (English)

Recent successes in self-supervised learning (SSL) model spatial co-occurrences of visual features either by masking portions of an image or by aggressively cropping it. Here, we propose a new way to model spatial co-occurrences by aligning local representations (before pooling) with a global image representation. We present CO-SSL, a family of instance discrimination methods and show that it outperforms previous methods on several datasets, including ImageNet-1K where it achieves 71.5% of Top-1 accuracy with 100 pre-training epochs. CO-SSL is also more robust to noise corruption, internal corruption, small adversarial attacks, and large training crop sizes. Our analysis further indicates that CO-SSL learns highly redundant local representations, which offers an explanation for its robustness. Overall, our work suggests that aligning local and global representations may be a powerful principle of unsupervised category learning.

自监督学习特征对齐图像表征鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。