arXiv:2505.24134stat.MLcs.CV2025-05被引 3

从概率视角重看对比学习,揭示其在多模态对齐中的数学本质。

A Mathematical Perspective On Contrastive Learning

论文配图:A Mathematical Perspective On Contrastive Learning
图 1 · 摘自论文原文
  • 将对比学习建模为条件概率分布的优化问题,实现跨模态对齐。
  • 在高斯设定下证明可逼近条件均值与协方差,支持生成与检索任务。
  • 适用于多模态检索、分类及生成,尤其适合海洋学等数据融合场景。

多模态对比学习是一种关联不同数据模态的方法,典型例子是图像与文本的对齐。本文聚焦双模态情形,将对比学习视为优化参数化编码器以定义条件概率分布(每个模态关于另一模态的条件分布),并使其与实际数据一致。该框架支持跨模态检索(识别条件分布的众数)与跨模态分类(带微调的任务适配)。同时衍生出跨模态生成模型。该概率视角引出两类自然推广:新型概率损失函数,以及用于共同隐空间对齐的替代度量。我们在多元高斯设定下研究这些推广,将隐空间识别视为低秩矩阵逼近问题。这使得我们能刻画损失函数与对齐度量对自然统计量(如条件均值、协方差)的逼近能力,从而提出针对模式寻找与生成任务的新算法。框架在多元高斯、标注MNIST数据集及海洋学数据同化应用中通过数值实验验证。

原文摘要 · Abstract (English)

Multimodal contrastive learning is a methodology for linking different data modalities; the canonical example is linking image and text data. The methodology is typically framed as the identification of a set of encoders, one for each modality, that align representations within a common latent space. In this work, we focus on the bimodal setting and interpret contrastive learning as the optimization of (parameterized) encoders that define conditional probability distributions, for each modality conditioned on the other, consistent with the available data. This provides a framework for multimodal algorithms such as crossmodal retrieval, which identifies the mode of one of these conditional distributions, and crossmodal classification, which is similar to retrieval but includes a fine-tuning step to make it task specific. The framework we adopt also gives rise to crossmodal generative models. This probabilistic perspective suggests two natural generalizations of contrastive learning: the introduction of novel probabilistic loss functions, and the use of alternative metrics for measuring alignment in the common latent space. We study these generalizations of the classical approach in the multivariate Gaussian setting. In this context we view the latent space identification as a low-rank matrix approximation problem. This allows us to characterize the capabilities of loss functions and alignment metrics to approximate natural statistics, such as conditional means and covariances; doing so yields novel variants on contrastive learning algorithms for specific mode-seeking and for generative tasks. The framework we introduce is also studied through numerical experiments on multivariate Gaussians, the labeled MNIST dataset, and on a data assimilation application arising in oceanography.

对比学习概率建模多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。