arXiv:2507.21395cs.MMcs.AI2025-07被引 1

用图注意力建模多模态情感,提升跨模态融合效果

Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion

  • 为每种模态设计动态增强模块,构建异构跨模态图
  • 在MELD和IEMOCAP上准确率与加权F1均超越当前最佳
  • 适合做多模态情感分析、特别是数据不均衡场景

多模态情感识别对实现具有情感智能的系统至关重要。然而,现有方法存在跨模态交互有限及各模态贡献不均的问题。为此,我们提出Sync-TVA,一种端到端的图注意力框架,包含模态特定的动态增强与结构化跨模态融合。该框架为每种模态设计动态增强模块,并构建异构跨模态图以建模文本、音频和视觉特征间的语义关系。通过跨注意力融合机制进一步对齐多模态线索,实现鲁棒的情感推断。在MELD和IEMOCAP数据集上的实验表明,该方法在准确率和加权F1分数上均持续优于现有最优模型,尤其在类别不平衡条件下表现突出。

原文摘要 · Abstract (English)

Multimodal emotion recognition (MER) is crucial for enabling emotionally intelligent systems that perceive and respond to human emotions. However, existing methods suffer from limited cross-modal interaction and imbalanced contributions across modalities. To address these issues, we propose Sync-TVA, an end-to-end graph-attention framework featuring modality-specific dynamic enhancement and structured cross-modal fusion. Our design incorporates a dynamic enhancement module for each modality and constructs heterogeneous cross-modal graphs to model semantic relations across text, audio, and visual features. A cross-attention fusion mechanism further aligns multimodal cues for robust emotion inference. Experiments on MELD and IEMOCAP demonstrate consistent improvements over state-of-the-art models in both accuracy and weighted F1 score, especially under class-imbalanced conditions.

情感识别多模态图神经网络跨模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。