解释对比学习为何让多模态模型通用,揭示其理论优势。
A Statistical Theory of Contrastive Pre-training and Multimodal Generative AI
- 用近似充分统计量解释对比预训练的有效性
- 证明预训练模型在零样本分类等任务上表现优异
- 适合研究多模态学习理论或模型泛化性的学者
多模态生成式AI系统(如视觉-语言模型)依赖对比预训练来学习跨模态表征。尽管实践效果显著,但其理论机制仍不清晰。本文提出近似充分统计量概念,证明对比预训练损失的近似最小值解具有近似充分性,可适配多种下游任务。构建图像与文本联合生成分层模型,证明变换器可通过信念传播高效逼近该模型中的关键函数。基于此框架,推导出基于对比预训练表示的多模态学习样本复杂度边界。数值实验验证了理论结果,表明对比预训练变换器在各类多模态任务中具备强大泛化能力。
原文摘要 · Abstract (English)
Multi-modal generative AI systems, such as those combining vision and language, rely on contrastive pre-training to learn representations across different modalities. While their practical benefits are widely acknowledged, a rigorous theoretical understanding of the contrastive pre-training framework remains limited. This paper develops a theoretical framework to explain the success of contrastive pre-training in downstream tasks, such as zero-shot classification, conditional diffusion models, and vision-language models. We introduce the concept of approximate sufficient statistics, a generalization of the classical sufficient statistics, and show that near-minimizers of the contrastive pre-training loss are approximately sufficient, making them adaptable to diverse downstream tasks. We further propose the Joint Generative Hierarchical Model for the joint distribution of images and text, showing that transformers can efficiently approximate relevant functions within this model via belief propagation. Building on this framework, we derive sample complexity guarantees for multi-modal learning based on contrastive pre-trained representations. Numerical simulations validate these theoretical findings, demonstrating the strong generalization performance of contrastively pre-trained transformers in various multi-modal tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。