arXiv:2509.24734cs.LGcs.AI2025-09NeurIPS被引 15

用三角形面积替代余弦相似度,提升三模态对齐效果

A TRIANGLE Enables Multimodal Alignment Beyond Cosine Similarity

  • 基于模态嵌入的高维空间计算三角形面积作为相似度
  • 在视频-文本、音频-文本检索中提升Recall@1达9个百分点
  • 无需额外融合层,结果可解释,适合多模态模型优化

多模态学习通过融合多种模态信息构建更全面的表征,在推动人工智能发展方面至关重要。然而,当前先进模型仍存在严重局限,难以确保所有模态有效对齐。部分模态可能未对齐,导致下游任务中无法充分挖掘多模态带来的信息增益。本文提出TRIANGLE:一种在模态嵌入构成的高维空间中直接计算的新型相似度度量,通过三角形面积实现三模态联合对齐,避免使用额外融合层或成对相似度。将TRIANGLE替换对比损失中的余弦相似度后,显著提升多模态建模性能,并提供可解释的对齐依据。在视频-文本、音频-文本检索及音频-视频分类等三模态任务中,TRIANGLE在多个数据集上达到领先水平,相较基于余弦相似度的方法,Recall@1最高提升9个百分点。

原文摘要 · Abstract (English)

Multimodal learning plays a pivotal role in advancing artificial intelligence systems by incorporating information from multiple modalities to build a more comprehensive representation. Despite its importance, current state-of-the-art models still suffer from severe limitations that prevent the successful development of a fully multimodal model. Such methods may not provide indicators that all the involved modalities are effectively aligned. As a result, some modalities may not be aligned, undermining the effectiveness of the model in downstream tasks where multiple modalities should provide additional information that the model fails to exploit. In this paper, we present TRIANGLE: TRI-modAl Neural Geometric LEarning, the novel proposed similarity measure that is directly computed in the higher-dimensional space spanned by the modality embeddings. TRIANGLE improves the joint alignment of three modalities via a triangle-area similarity, avoiding additional fusion layers or pairwise similarities. When incorporated in contrastive losses replacing cosine similarity, TRIANGLE significantly boosts the performance of multimodal modeling, while yielding interpretable alignment rationales. Extensive evaluation in three-modal tasks such as video-text and audio-text retrieval or audio-video classification, demonstrates that TRIANGLE achieves state-of-the-art results across different datasets improving the performance of cosine-based methods up to 9 points of Recall@1.

多模态相似度度量对齐优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。