arXiv:2503.22715cs.LGcs.CV2025-03被引 3

提出分层自适应专家框架,提升多模态情感分析的跨模态融合能力。

Hierarchical Adaptive Expert for Multimodal Sentiment Analysis

  • 设计分层自适应专家结构,分别捕捉全局与局部模态特征。
  • 在CMU-MOSEI等数据集上准确率提升2.6%~6.3%,误差降低0.058~0.059。
  • 适用于部分或完整模态输入场景,适合需高精度情感理解的任务。

多模态情感分析已成为理解多元沟通渠道中人类情绪的关键工具。现有方法常难以有效区分和融合共享与特异的模态信息,限制了多模态学习性能。为此,本文提出分层自适应专家多模态情感分析框架(HAEMSA),融合进化优化、跨模态知识迁移与多任务学习。HAEMSA采用分层自适应专家结构,捕获全局与局部模态表征,实现更精细的情感分析。通过进化算法动态优化网络结构与模态组合,适应部分或全模态场景。大量实验表明,HAEMSA在多个基准数据集上表现优异:在CMU-MOSEI上,7类准确率提升2.6%,平均绝对误差(MAE)降低0.059;在CMU-MOSI上,7类准确率提高6.3%,MAE下降0.058;在IEMOCAP上,加权F1得分超越当前最佳方法2.84%。结果证明该方法能有效捕捉复杂多模态交互,并在不同情感语境中良好泛化。

原文摘要 · Abstract (English)

Multimodal sentiment analysis has emerged as a critical tool for understanding human emotions across diverse communication channels. While existing methods have made significant strides, they often struggle to effectively differentiate and integrate modality-shared and modality-specific information, limiting the performance of multimodal learning. To address this challenge, we propose the Hierarchical Adaptive Expert for Multimodal Sentiment Analysis (HAEMSA), a novel framework that synergistically combines evolutionary optimization, cross-modal knowledge transfer, and multi-task learning. HAEMSA employs a hierarchical structure of adaptive experts to capture both global and local modality representations, enabling more nuanced sentiment analysis. Our approach leverages evolutionary algorithms to dynamically optimize network architectures and modality combinations, adapting to both partial and full modality scenarios. Extensive experiments demonstrate HAEMSA's superior performance across multiple benchmark datasets. On CMU-MOSEI, HAEMSA achieves a 2.6% increase in 7-class accuracy and a 0.059 decrease in MAE compared to the previous best method. For CMU-MOSI, we observe a 6.3% improvement in 7-class accuracy and a 0.058 reduction in MAE. On IEMOCAP, HAEMSA outperforms the state-of-the-art by 2.84% in weighted-F1 score for emotion recognition. These results underscore HAEMSA's effectiveness in capturing complex multimodal interactions and generalizing across different emotional contexts.

情感分析多模态自适应进化优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。