arXiv:2607.17673cs.LGcs.AI2026-07

改善编码器几何性质可显著提升多模态对比学习效果

Beyond Objective Expressivity: Geometry Preservation in Multimodal Contrastive Learning

论文配图:Beyond Objective Expressivity: Geometry Preservation in Multimodal Contrastive Learning
图 1 · 摘自论文原文
  • 通过正则化控制编码器雅可比条件数,防止特征退化
  • 在4个真实数据集上,检索与线性探测性能均明显提升
  • 适用于存在缺失模态的复杂场景,对模型设计有指导意义

对比学习正从图像-文本配对拓展至三模态及以上。然而,高阶对齐易引发优化与表征难题。我们发现编码器雅可比条件数是三模态对比学习的关键:条件不佳会导致奇异值谱崩溃或放大,引发雅可比条件数爆炸,损害多模态对齐。本文提出几何保真编码器(GPE),通过正则化直接调控雅可比条件数,并验证了LeakyReLU激活和残差路径等简单改进即可恢复几何优势。在合成基准与四个真实数据集(含缺失模态)上,优化雅可比条件数显著提升检索与线性探测性能,而仅提升目标表达能力则效果有限。结果表明,多模态对比学习不仅依赖目标表达力,更受编码器几何与优化特性的制约。

原文摘要 · Abstract (English)

Contrastive learning is increasingly moving toward settings with three or more modalities instead of image-text pairs. Yet, extending models from pairwise to higher-order multimodal alignment can introduce optimization and representation challenges. We identify encoder Jacobian conditioning as a key factor in trimodal contrastive learning: poorly conditioned encoders exhibit collapsing or amplified singular-value spectra, leading to exploding Jacobian condition numbers and degraded multimodal alignment. We introduce geometry-preserving encoders (GPEs) by directly conditioning the Jacobian through regularization and demonstrating that simple modifications like LeakyReLU activations and residual paths recover these geometric benefits. Across a synthetic benchmark and four real-world datasets including missing modalities, improving Jacobian conditioning boosts retrieval and linear probe performance across multiple contrastive objectives, whereas expressive objectives yield little benefit in linear probes. More broadly, our results show that multimodal contrastive learning depends not only on objective expressivity, but also on the geometric and optimization properties of the underlying encoders.

多模态学习对比学习几何建模编码器设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。