arXiv:2508.18322cs.CVcs.AI2025-08AAAI被引 1

通过图对比学习融合多模态结构与语义,提升情感分析性能与可解释性。

Structures Meet Semantics: Multimodal Fusion via Graph Contrastive Learning

  • 构建模态专属图结构,用语法和轻量注意力捕捉跨模态关系。
  • 引入全局文本语义锚点,对齐异构模态的语义空间,提升一致性。
  • 多视角对比学习增强表示判别力,适合需要可解释性的场景。

多模态情感分析(MSA)旨在通过有效整合文本、语音和视觉模态来推断情绪状态。尽管取得显著进展,现有融合方法常忽略模态特异性结构依赖和语义错位,限制了其质量、可解释性和鲁棒性。为此,我们提出一种新框架——结构-语义统一器(SSU),系统整合模态特异性结构信息与跨模态语义对齐,以增强多模态表征。具体而言,SSU通过语言句法构建文本图,利用轻量级文本引导注意力机制构建语音与视觉图,从而捕捉细粒度的模态内关系与语义交互。我们进一步引入一个源自全局文本语义的语义锚点,作为跨模态对齐枢纽,有效调和不同模态间的异质语义空间。此外,设计多视图对比学习目标,促进模态内与跨模态视图间的可区分性、语义一致性和结构连贯性。在两个广泛使用的基准数据集CMU-MOSI和CMU-MOSEI上的大量实验表明,SSU持续实现最先进性能,同时相比以往方法显著降低计算开销。全面的定性分析进一步验证了其可解释性及通过语义驱动交互捕捉细微情感模式的能力。

原文摘要 · Abstract (English)

Multimodal sentiment analysis (MSA) aims to infer emotional states by effectively integrating textual, acoustic, and visual modalities. Despite notable progress, existing multimodal fusion methods often neglect modality-specific structural dependencies and semantic misalignment, limiting their quality, interpretability, and robustness. To address these challenges, we propose a novel framework called the Structural-Semantic Unifier (SSU), which systematically integrates modality-specific structural information and cross-modal semantic grounding for enhanced multimodal representations. Specifically, SSU dynamically constructs modality-specific graphs by leveraging linguistic syntax for text and a lightweight, text-guided attention mechanism for acoustic and visual modalities, thus capturing detailed intra-modal relationships and semantic interactions. We further introduce a semantic anchor, derived from global textual semantics, that serves as a cross-modal alignment hub, effectively harmonizing heterogeneous semantic spaces across modalities. Additionally, we develop a multiview contrastive learning objective that promotes discriminability, semantic consistency, and structural coherence across intra- and inter-modal views. Extensive evaluations on two widely used benchmark datasets, CMU-MOSI and CMU-MOSEI, demonstrate that SSU consistently achieves state-of-the-art performance while significantly reducing computational overhead compared to prior methods. Comprehensive qualitative analyses further validate SSU's interpretability and its ability to capture nuanced emotional patterns through semantically grounded interactions.

多模态情感分析图神经网络对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。