通过对齐局部与全局特征,提升自监督学习的性能与鲁棒性。
Seeing the Whole in the Parts in Self-Supervised Representation Learning
- 用局部表示与全局图像表示对齐建模空间关系
- ImageNet-1K上达71.5%准确率,优于此前方法
- 对噪声、对抗攻击和大裁剪更鲁棒,适合工业级应用
自监督学习(SSL)近期通过遮蔽图像部分或大幅裁剪来建模视觉特征的空间共现。本文提出一种新方法——CO-SSL,通过将局部表示(池化前)与全局图像表示对齐来建模空间共现。该方法为实例判别类算法家族,在多个数据集上表现优于以往方法,尤其在ImageNet-1K上以100个预训练周期达到71.5%的Top-1准确率。此外,CO-SSL在面对噪声污染、内部干扰、小规模对抗攻击及大训练裁剪尺寸时均表现出更强鲁棒性。分析表明,其学习到的高度冗余局部表示可能是鲁棒性的来源。整体上,本研究提示:对齐局部与全局表示或可成为无监督类别学习的关键原则。
原文摘要 · Abstract (English)
Recent successes in self-supervised learning (SSL) model spatial co-occurrences of visual features either by masking portions of an image or by aggressively cropping it. Here, we propose a new way to model spatial co-occurrences by aligning local representations (before pooling) with a global image representation. We present CO-SSL, a family of instance discrimination methods and show that it outperforms previous methods on several datasets, including ImageNet-1K where it achieves 71.5% of Top-1 accuracy with 100 pre-training epochs. CO-SSL is also more robust to noise corruption, internal corruption, small adversarial attacks, and large training crop sizes. Our analysis further indicates that CO-SSL learns highly redundant local representations, which offers an explanation for its robustness. Overall, our work suggests that aligning local and global representations may be a powerful principle of unsupervised category learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。