让视觉模型更省力、更懂意思,通过稀疏化提升表征质量
SparseJEPA: Sparse Representation Learning of Joint Embedding Predictive Architectures
- 引入稀疏性约束,让语义相关的特征共享潜在变量
- 在CIFAR-100上训练的轻量ViT实现更优线性探针迁移性能
- 理论证明分组机制降低潜变量间多信息量,提升可解释性
联合嵌入预测架构(JEPA)已成为学习通用表征的强大框架,但其密集嵌入表示常导致可解释性差与效率低。本文提出SparseJEPA,将稀疏表征学习融入JEPA框架,通过惩罚项促使语义相关特征共享潜在变量,同时保持预测性能。我们在CIFAR-100上训练轻量级视觉变换器,并用改进的嵌入进行线性探针迁移学习,涵盖图像分类与低层任务,展现架构在不同迁移任务中的泛化能力。进一步地,我们提供了理论证明:分组机制能减少潜变量间的多信息量,包括对多信息量数据处理不等式的证明。结果表明,引入稀疏性不仅优化了潜在空间,还促进了更有意义且可解释的表征学习。未来工作将探索通过以对象为中心的表征学习,进一步挖掘该机制潜力。
原文摘要 · Abstract (English)
Joint Embedding Predictive Architectures (JEPA) have emerged as a powerful framework for learning general-purpose representations. However, these models often lack interpretability and suffer from inefficiencies due to dense embedding representations. We propose SparseJEPA, an extension that integrates sparse representation learning into the JEPA framework to enhance the quality of learned representations. SparseJEPA employs a penalty method that encourages latent space variables to be shared among data features with strong semantic relationships, while maintaining predictive performance. We demonstrate the effectiveness of SparseJEPA by training on the CIFAR-100 dataset and pre-training a lightweight Vision Transformer. The improved embeddings are utilized in linear-probe transfer learning for both image classification and low-level tasks, showcasing the architecture's versatility across different transfer tasks. Furthermore, we provide a theoretical proof that demonstrates that the grouping mechanism enhances representation quality. This was done by displaying that grouping reduces Multiinformation among latent-variables, including proofing the Data Processing Inequality for Multiinformation. Our results indicate that incorporating sparsity not only refines the latent space but also facilitates the learning of more meaningful and interpretable representations. In further work, hope to further extend this method by finding new ways to leverage the grouping mechanism through object-centric representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。