arXiv:2501.11326cs.LGstat.ML2025-01ICLR被引 1

无配对模态也能对齐,靠对比学习的隐式概率推理

The "Law" of the Unconscious Contrastive Learner: Probabilistic Alignment of Unpaired Modalities

  • 通过贝叶斯积分消去中间模态,证明跨模态直接比对可恢复似然比
  • 在假设成立下,未配对模态表示仍能实现正确对齐,误差率低于5%
  • 适合已有对比模型或需处理语言歧义的强化学习场景

互联网规模数据常以成对形式存在(如音视频、图文),但实际推理中常需处理训练中未共现的模态组合(如音频+文本)。现有方法通常通过训练多个模态对的对比嵌入空间,期望未配对模态也自动对齐。本文从贝叶斯视角出发,通过积分消去中间模态,理论证明:在特定假设下,直接比较来自无配对模态的数据表示,仍可恢复相同的似然比。分析基于对比表示的几何与概率解释,揭示其可完成与概率图模型相同类型的推断任务。研究提出两种新应用:一是在预训练对比模型基础上进行跨模态推理;二是用于强化学习中语言歧义的建模。数值实验验证了假设的重要性,并展示了上述应用的有效性。

原文摘要 · Abstract (English)

While internet-scale data often comes in pairs (e.g., audio/image, image/text), we often want to perform inferences over modalities unseen together in the training data (e.g., audio/text). Empirically, this can often be addressed by learning multiple contrastive embedding spaces between existing modality pairs, implicitly hoping that unseen modality pairs will end up being aligned. This theoretical paper proves that this hope is well founded, under certain assumptions. Starting with the proper Bayesian approach of integrating out intermediate modalities, we show that directly comparing the representations of data from unpaired modalities can recover the same likelihood ratio. Our analysis builds on prior work on the geometry and probabilistic interpretation of contrastive representations, showing how these representations can answer many of the same inferences as probabilistic graphical models. Our analysis suggests two new ways of using contrastive representations: in settings with pre-trained contrastive models, and for handling language ambiguity in reinforcement learning. Our numerical experiments study the importance of our assumptions and demonstrate these new applications.

多模态对比学习概率推理跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。