对比多模态与单模态对比学习,发现信号噪声比是性能关键
On the Comparison between Multi-modal and Single-modal Contrastive Learning
- 基于信号噪声模型分析对比学习优化过程
- 多模态通过模态协作提升信号噪声比,增强泛化能力
- 理论框架适用于单/多模态,适合研究表示学习机制者
多模态对比学习在语言监督下已成为现代机器学习的新范式。通过在大规模网络数据上预训练,其可学习高质量表征,展现出出色的鲁棒性与迁移能力。尽管实证成功,其理论理解仍不充分,尤其缺乏与单模态对比学习的比较。本文提出一个特征学习理论框架,基于包含信号与噪声的数据生成模型,对使用InfoMax目标函数的ReLU网络进行轨迹优化分析与下游任务泛化刻画。结果表明,信号-噪声比(SNR)是影响单/多模态对比学习下游泛化能力的关键因素。多模态通过双模态协作提升了有效信号,从而在下游任务中表现优于单模态。该分析建立了统一框架,能刻画单/多模态对比学习的优化与泛化特性。合成与真实数据集上的实验进一步验证了理论结论。
原文摘要 · Abstract (English)
Multi-modal contrastive learning with language supervision has presented a paradigm shift in modern machine learning. By pre-training on a web-scale dataset, multi-modal contrastive learning can learn high-quality representations that exhibit impressive robustness and transferability. Despite its empirical success, the theoretical understanding is still in its infancy, especially regarding its comparison with single-modal contrastive learning. In this work, we introduce a feature learning theory framework that provides a theoretical foundation for understanding the differences between multi-modal and single-modal contrastive learning. Based on a data generation model consisting of signal and noise, our analysis is performed on a ReLU network trained with the InfoMax objective function. Through a trajectory-based optimization analysis and generalization characterization on downstream tasks, we identify the critical factor, which is the signal-to-noise ratio (SNR), that impacts the generalizability in downstream tasks of both multi-modal and single-modal contrastive learning. Through the cooperation between the two modalities, multi-modal learning can achieve better feature learning, leading to improvements in performance in downstream tasks compared to single-modal learning. Our analysis provides a unified framework that can characterize the optimization and generalization of both single-modal and multi-modal contrastive learning. Empirical experiments on both synthetic and real-world datasets further consolidate our theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。