arXiv:2511.12121cs.LGcs.MM2025-11AAAI被引 22

明确对齐未必更好,需根据模态冗余度调整对齐强度

To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance

  • 设计可控对比学习模块,可调节不同模态间表示对齐程度
  • 实验证明:对齐过强或过弱都会降低单模态模型性能
  • 发现最优对齐强度取决于模态间信息冗余度,适合调参参考

多模态学习通常依赖跨模态表示对齐以实现有效信息融合,这一做法常被视为普遍有益。然而,以往研究多为观察性分析,仅考察自然数据中对齐与模型性能的相关性,未系统探究显式强制对齐对表示的影响。本文引入一个可控的对比学习模块,可在训练中精确调控对齐强度,从而探索在不同模态信息结构下,显式对齐如何影响模型性能与表示对齐。在合成与真实数据集上的实验表明,显式对齐对单模态模型性能的影响取决于数据特征:最优对齐强度由模态间冗余量决定。我们识别出一个平衡模态特异性信号与共享冗余信息的最优对齐水平,为提升单模态编码器性能提供了实用指导。

原文摘要 · Abstract (English)

Multimodal learning often relies on aligning representations across modalities to enable effective information integration, an approach traditionally assumed to be universally beneficial. However, prior research has primarily taken an observational approach, examining naturally occurring alignment in multimodal data and exploring its correlation with model performance, without systematically studying the direct effects of explicitly enforced alignment between representations of different modalities. In this work, we investigate how explicit alignment influences both model performance and representation alignment under different modality-specific information structures. Specifically, we introduce a controllable contrastive learning module that enables precise manipulation of alignment strength during training, allowing us to explore when explicit alignment improves or hinders performance. Our results on synthetic and real datasets under different data characteristics show that the impact of explicit alignment on the performance of unimodal models is related to the characteristics of the data: the optimal level of alignment depends on the amount of redundancy between the different modalities. We identify an optimal alignment strength that balances modality-specific signals and shared redundancy in the mixed information distributions. This work provides practical guidance on when and how explicit alignment should be applied to achieve optimal unimodal encoder performance.

多模态表示对齐对比学习模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。