通过引入隐变量解决视觉自监督学习中的多目标映射问题
Self-Supervised Learning from Structural Invariance
- 用隐变量建模数据对间的条件不确定性,提升灵活性
- 在视频因果推理、细粒度图像理解等任务上性能显著提升
- 适用于对比学习与蒸馏类方法,通用性强
联合嵌入自监督学习(SSL)是视觉数据无监督表征学习的核心范式,其通过语义相关数据对之间的不变性进行学习。我们研究了SSL中的“一对多”映射问题,即每个样本可能对应多个有效目标,这在自然生成过程中(如连续视频帧)普遍存在。现有方法难以灵活捕捉这种条件不确定性。为此,我们引入隐变量以建模该不确定性,并推导出配对嵌入间互信息的变分下界。该推导得出一个适用于标准SSL目标的简单正则化项。由此提出的AdaSSL方法可应用于对比学习和基于蒸馏的SSL目标,并在因果表征学习、细粒度图像理解及视频世界建模任务中展现出良好泛化能力。
原文摘要 · Abstract (English)
Joint-embedding self-supervised learning (SSL), the key paradigm for unsupervised representation learning from visual data, learns from invariances between semantically-related data pairs. We study the one-to-many mapping problem in SSL, where each datum may be mapped to multiple valid targets. This arises when data pairs come from naturally occurring generative processes, e.g., successive video frames. We show that existing methods struggle to flexibly capture this conditional uncertainty. As a remedy, we introduce a latent variable to account for this uncertainty and derive a variational lower bound on the mutual information between paired embeddings. Our derivation yields a simple regularization term for standard SSL objectives. The resulting method, which we call AdaSSL, applies to both contrastive and distillation-based SSL objectives, and we empirically show its versatility in causal representation learning, fine-grained image understanding, and world modeling on videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。