arXiv:2602.05639cs.LGstat.ML2026-02

提出一种无需重建的自监督表示学习框架,直接在嵌入空间建模不确定性。

Joint Embedding Variational Bayes

  • 用球面分布定义目标嵌入的条件似然,分离方向与幅度一致性
  • 在ImageNet等数据集上性能媲美主流非对比方法,支持下游不确定性感知任务
  • 无需额外投影头,实现特征级异方差不确定性建模,适合鲁棒性要求高的场景

我们提出变分联合嵌入(VJE),一种用于表示空间非对比自监督学习的无重建潜变量框架。VJE通过在配对编码器嵌入上最大化对称条件证据下界(ELBO),直接在目标表示上定义条件似然,而非优化点对兼容性目标。似然以目标嵌入的极坐标表示上的重尾Student-t分布实现,其方向-径向分解可分离方向一致性与幅度一致性,并缓解范数诱导的病态问题。方向分量作用于单位球面,为对应的球面子密度模型提供有效变分界。一个摊销推断网络参数化对角高斯后验,其特征级方差与方向似然共享,实现无辅助投影头的各向异性不确定性建模。在ImageNet-1K、CIFAR-10/100和STL-10上,VJE在线性及k-NN评估中表现与标准非对比基线相当,且在表示空间中直接提供概率语义,适用于下游不确定性感知应用。通过分布外检测验证,表示空间中的似然值表现出强实际性能。结果表明该框架是非对比学习的一种原则性变分形式,可在学习嵌入空间中直接表示结构化特征级不确定性。

原文摘要 · Abstract (English)

We introduce Variational Joint Embedding (VJE), a reconstruction-free latent-variable framework for non-contrastive self-supervised learning in representation space. VJE maximizes a symmetric conditional evidence lower bound (ELBO) on paired encoder embeddings by defining a conditional likelihood directly on target representations, rather than optimizing a pointwise compatibility objective. The likelihood is instantiated as a heavy-tailed Student--\(t\) distribution on a polar representation of the target embedding, where a directional--radial decomposition separates angular agreement from magnitude consistency and mitigates norm-induced pathologies. The directional factor operates on the unit sphere, yielding a valid variational bound for the associated spherical subdensity model. An amortized inference network parameterizes a diagonal Gaussian posterior whose feature-wise variances are shared with the directional likelihood, yielding anisotropic uncertainty without auxiliary projection heads. Across ImageNet-1K, CIFAR-10/100, and STL-10, VJE is competitive with standard non-contrastive baselines under linear and \(k\)-NN evaluation, while providing probabilistic semantics directly in representation space for downstream uncertainty-aware applications. We validate these semantics through out-of-distribution detection, where representation-space likelihoods yield strong empirical performance. These results position the framework as a principled variational formulation of non-contrastive learning, in which structured feature-wise uncertainty is represented directly in the learned embedding space.

自监督学习不确定性建模嵌入空间变分推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。