arXiv:2508.16300cs.CVcs.AI2025-08被引 8

不直接交互模态,用关系图与分层注意力提升多模态理解

A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension

  • 通过跨模态关系图重构特征,避免模态间直接交互引入噪声
  • 在三个数据集上实现多任务性能提升,关键指标优于基线模型
  • 适合需要保留单模态判别信息的复杂多模态应用

多模态学习中的主要挑战是各模态内部存在噪声,这些噪声会通过显式模态交互影响最终的融合表示。现有融合方法虽旨在获得强联合表示,却可能忽略单个模态中的有价值判别信息。为此,我们提出一种多模态多任务框架——MM-ORIENT,其通过跨模态关系图在不进行显式交互的前提下获取多模态表示,从而降低潜在空间中的噪声影响。该方法基于不同模态特征决定节点邻域,重构单模态特征以形成多模态表示。同时,我们设计了分层交互单模态注意力(HIMA),聚焦于模态内部相关特征。跨模态关系图用于捕捉两模态间的高阶关联,而HIMA则在晚期融合前学习各模态的判别特征,支持多任务学习。在三个数据集上的大量实验表明,该方法能有效实现多模态内容的语义理解,适用于多种任务。

原文摘要 · Abstract (English)

A major challenge in multimodal learning is the presence of noise within individual modalities. This noise inherently affects the resulting multimodal representations, especially when these representations are obtained through explicit interactions between different modalities. Moreover, the multimodal fusion techniques while aiming to achieve a strong joint representation, can neglect valuable discriminative information within the individual modalities. To this end, we propose a Multimodal-Multitask framework with crOss-modal Relation and hIErarchical iNteractive aTtention (MM-ORIENT) that is effective for multiple tasks. The proposed approach acquires multimodal representations cross-modally without explicit interaction between different modalities, reducing the noise effect at the latent stage. To achieve this, we propose cross-modal relation graphs that reconstruct monomodal features to acquire multimodal representations. The features are reconstructed based on the node neighborhood, where the neighborhood is decided by the features of a different modality. We also propose Hierarchical Interactive Monomadal Attention (HIMA) to focus on pertinent information within a modality. While cross-modal relation graphs help comprehend high-order relationships between two modalities, HIMA helps in multitasking by learning discriminative features of individual modalities before late-fusing them. Finally, extensive experimental evaluation on three datasets demonstrates that the proposed approach effectively comprehends multimodal content for multiple tasks.

多模态注意力机制特征融合多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。