对比学习能自动发现数据内在维度,学出更紧凑的有用表示。
Multi-modal contrastive learning adapts to intrinsic dimensions of shared latent variables
- 通过温度优化,让多模态对比学习自动适应数据真实低维结构。
- 在合成与真实数据集上均学出低维且信息丰富的表示,维度远低于预设值。
- 理论与实验结合,适合研究表示学习机制或模型压缩的读者。
多模态对比学习作为自监督表征学习技术,在基础模型训练中取得显著成功,如CLIP。本文研究了多模态对比学习在非线性表征和特定数据分布之外的理论性质。分析表明,借助温度优化,该方法不仅能最大化模态间互信息,还能自适应数据的内在维度,该维度远低于用户指定的表征向量维度。在合成与真实世界数据集上的实验验证了对比学习学习低维、高信息量表示的能力,连接了理论洞见与实际性能。
原文摘要 · Abstract (English)
Multi-modal contrastive learning as a self-supervised representation learning technique has achieved great success in foundation model training, such as CLIP~\citep{radford2021learning}. In this paper, we study the theoretical properties of the learned representations from multi-modal contrastive learning beyond linear representations and specific data distributions. Our analysis reveals that, enabled by temperature optimization, multi-modal contrastive learning not only maximizes mutual information between modalities but also adapts to intrinsic dimensions of data, which can be much lower than user-specified dimensions for representation vectors. Experiments on both synthetic and real-world datasets demonstrate the ability of contrastive learning to learn low-dimensional and informative representations, bridging theoretical insights and practical performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。