用概率模型替代确定性预测,提升自监督学习在多模态任务中的表现。
Gaussian Joint Embeddings For Self-Supervised Representation Learning
- 基于高斯联合嵌入建模上下文与目标的联合分布,实现闭式条件推断。
- 在合成数据和视觉基准上,显著恢复复杂条件结构并生成更优无条件样本。
- 提出多种抗崩溃机制,适用于多模态、非参数及动态结构场景。
自监督表示学习通常依赖确定性预测架构来对齐潜在空间中的上下文与目标视图。尽管在许多场景下有效,这类方法在真正的多模态逆问题中受限于平方损失预测趋向条件均值,且常依赖架构不对称性防止表征崩溃。本文提出基于生成联合建模的概率替代方案,引入高斯联合嵌入(GJE)及其多模态扩展——高斯混合联合嵌入(GMJE),建模上下文与目标表示的联合密度,以显式概率模型下的闭式条件推断替代黑箱预测。该方法提供合理的不确定性估计和协方差感知的目标函数以控制潜在几何结构。我们进一步识别出朴素经验批优化中的失败模式——马哈拉诺比迹陷阱,并提出多种修复策略,涵盖参数化、自适应和非参数设置,包括基于原型的GMJE、条件混合密度网络(GMJE-MDN)、拓扑自适应生长神经气泡(GMJE-GNG)以及序列蒙特卡洛(SMC)记忆库。此外,我们证明标准对比学习可被解释为GMJE框架在退化非参数极限下的特例。在合成多模态对齐任务和视觉基准上的实验表明,GMJE能恢复复杂条件结构,学习具有竞争力的判别性表示,并构建更适合无条件采样的潜在密度,优于确定性或单峰基线。
原文摘要 · Abstract (English)
Self-supervised representation learning often relies on deterministic predictive architectures to align context and target views in latent space. While effective in many settings, such methods are limited in genuinely multi-modal inverse problems, where squared-loss prediction collapses towards conditional averages, and they frequently depend on architectural asymmetries to prevent representation collapse. In this work, we propose a probabilistic alternative based on generative joint modeling. We introduce Gaussian Joint Embeddings (GJE) and its multi-modal extension, Gaussian Mixture Joint Embeddings (GMJE), which model the joint density of context and target representations and replace black-box prediction with closed-form conditional inference under an explicit probabilistic model. This yields principled uncertainty estimates and a covariance-aware objective for controlling latent geometry. We further identify a failure mode of naive empirical batch optimization, which we term the Mahalanobis Trace Trap, and develop several remedies spanning parametric, adaptive, and non-parametric settings, including prototype-based GMJE, conditional Mixture Density Networks (GMJE-MDN), topology-adaptive Growing Neural Gas (GMJE-GNG), and a Sequential Monte Carlo (SMC) memory bank. In addition, we show that standard contrastive learning can be interpreted as a degenerate non-parametric limiting case of the GMJE framework. Experiments on synthetic multi-modal alignment tasks and vision benchmarks show that GMJE recovers complex conditional structure, learns competitive discriminative representations, and defines latent densities that are better suited to unconditional sampling than deterministic or unimodal baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。