改进JEPA模型,让无监督视觉学习更稳定高效
Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
- 用对比学习+协方差正则化替代原JEPA的预测机制
- 在ImageNet-1K上训练后线性探测与微调更快更准
- 适合做自监督视觉表征学习的研究者参考
近期无监督视觉表征学习中,联合嵌入预测架构(JEPA)通过创新的掩码策略从无标签图像中提取特征。尽管取得成功,仍存在两大问题:I-JEPA中的指数移动平均(EMA)无法防止完全坍缩,且其预测机制难以准确学习块表示的均值。为此,本文提出C-JEPA(对比-JEPA)框架,将基于图像的联合嵌入预测架构与方差-不变-协方差正则化(VICReg)策略结合,有效学习方差/协方差以防止完全坍缩,并确保增强视图下均值的不变性,从而克服上述局限。通过实证与理论分析,结果表明C-JEPA显著提升了视觉表征学习的稳定性与质量。在ImageNet-1K数据集上预训练后,其在线性探测和微调任务中表现出更快且更优的收敛性能。
原文摘要 · Abstract (English)
In recent advancements in unsupervised visual representation learning, the Joint-Embedding Predictive Architecture (JEPA) has emerged as a significant method for extracting visual features from unlabeled imagery through an innovative masking strategy. Despite its success, two primary limitations have been identified: the inefficacy of Exponential Moving Average (EMA) from I-JEPA in preventing entire collapse and the inadequacy of I-JEPA prediction in accurately learning the mean of patch representations. Addressing these challenges, this study introduces a novel framework, namely C-JEPA (Contrastive-JEPA), which integrates the Image-based Joint-Embedding Predictive Architecture with the Variance-Invariance-Covariance Regularization (VICReg) strategy. This integration is designed to effectively learn the variance/covariance for preventing entire collapse and ensuring invariance in the mean of augmented views, thereby overcoming the identified limitations. Through empirical and theoretical evaluations, our work demonstrates that C-JEPA significantly enhances the stability and quality of visual representation learning. When pre-trained on the ImageNet-1K dataset, C-JEPA exhibits rapid and improved convergence in both linear probing and fine-tuning performance metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。