arXiv:2509.11344cs.CVcs.LG2025-09被引 1

探索视图多样性对自监督学习的影响,发现适度差异能提升模型性能

Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning

  • 用视图差异度量替代传统一致性假设,改进自监督学习机制
  • 适度增加视图差异可提升分类与密集预测任务表现,过度则适得其反
  • 提出用地球移动距离(EMD)量化视图间信息相关性,指导框架设计

自监督学习(SSL)通常依赖实例一致性假设,即同一图像的不同视图应视为正样本对。然而,对于非标志性数据,不同视图可能包含不同物体或语义信息,该假设失效。本文通过大量消融实验表明,即使正样本对不满足严格实例一致性,SSL仍可学习有意义的表征。进一步分析发现,通过消除视图重叠或采用更小裁剪尺度增加视图多样性,可提升分类与密集预测任务的下游性能。但过度多样性会降低效果,表明存在最优视图差异范围。为此,我们引入地球移动距离(EMD)作为视图间互信息的估计器,发现中等EMD值与更好SSL学习效果相关,为未来框架设计提供洞见。我们在多种设置下验证了结论的鲁棒性与普适性。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) conventionally relies on the instance consistency paradigm, assuming that different views of the same image can be treated as positive pairs. However, this assumption breaks down for non-iconic data, where different views may contain distinct objects or semantic information. In this paper, we investigate the effectiveness of SSL when instance consistency is not guaranteed. Through extensive ablation studies, we demonstrate that SSL can still learn meaningful representations even when positive pairs lack strict instance consistency. Furthermore, our analysis further reveals that increasing view diversity, by enforcing zero overlapping or using smaller crop scales, can enhance downstream performance on classification and dense prediction tasks. However, excessive diversity is found to reduce effectiveness, suggesting an optimal range for view diversity. To quantify this, we adopt the Earth Mover's Distance (EMD) as an estimator to measure mutual information between views, finding that moderate EMD values correlate with improved SSL learning, providing insights for future SSL framework design. We validate our findings across a range of settings, highlighting their robustness and applicability on diverse data sources.

自监督学习视图多样性特征表示图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。