arXiv:2508.10133cs.CV2025-08NeurIPS被引 4

提出可解释的多模态融合模型,显式建模跨模态关联。

MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning

  • 设计可逆交叉注意力层,显式建模多模态间关系。
  • 在三个任务上达到当前最优性能,包括图像分割与电影分类。
  • 适合需要可解释性融合的多模态应用,如医疗分析。

多模态学习近年取得显著进展,但现有方法依赖Transformer的注意力机制隐式学习模态特征间的相关性,难以捕捉各模态的本质特征,导致对复杂结构和关联的理解不足。本文提出一种基于归一化流的新型多模态注意力融合方法(MANGO),通过引入可逆交叉注意力(ICA)层,实现显式、可解释且可计算的多模态融合。为高效捕获多模态数据中复杂的潜在关联,提出三种新交叉注意力机制:模态到模态交叉注意力(MMCA)、跨模态交叉注意力(IMCA)和可学习跨模态交叉注意力(LICA)。最后,构建多模态注意力归一化流以支持高维多模态数据的可扩展性。在语义分割、图像到图像转换及电影类型分类三个任务上的实验表明,该方法达到当前最优性能(SoTA)。

原文摘要 · Abstract (English)

Multimodal learning has gained much success in recent years. However, current multimodal fusion methods adopt the attention mechanism of Transformers to implicitly learn the underlying correlation of multimodal features. As a result, the multimodal model cannot capture the essential features of each modality, making it difficult to comprehend complex structures and correlations of multimodal inputs. This paper introduces a novel Multimodal Attention-based Normalizing Flow (MANGO) approach to developing explicit, interpretable, and tractable multimodal fusion learning. In particular, we propose a new Invertible Cross-Attention (ICA) layer to develop the Normalizing Flow-based Model for multimodal data. To efficiently capture the complex, underlying correlations in multimodal data in our proposed invertible cross-attention layer, we propose three new cross-attention mechanisms: Modality-to-Modality Cross-Attention (MMCA), Inter-Modality Cross-Attention (IMCA), and Learnable Inter-Modality Cross-Attention (LICA). Finally, we introduce a new Multimodal Attention-based Normalizing Flow to enable the scalability of our proposed method to high-dimensional multimodal data. Our experimental results on three different multimodal learning tasks, i.e., semantic segmentation, image-to-image translation, and movie genre classification, have illustrated the state-of-the-art (SoTA) performance of the proposed approach.

多模态融合归一化流可解释模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。