arXiv:2601.07565cs.CLcs.AI2026-01被引 1

用专家网络融合多模态数据,统一解决情绪识别与情感分析问题。

Expert-Guided Multimodal Fusion for Unified Emotion and Sentiment Analysis

  • 设计多尺度专家网络,分别捕捉细微情绪、跨模态关联和长程依赖。
  • 在多个双语数据集上实现更高准确率与跨语言鲁棒性。
  • 适合需要统一处理情绪与情感分析的研究者或开发者。

多模态情绪理解需融合文本、音频、视觉等异构数据,同时完成离散情绪识别与连续情感分析。我们提出EGMF框架,结合专家引导的多模态融合与大语言模型,在两项任务上均取得优越性能。核心为多尺度专家网络,包括局部专家(捕捉细微情绪)、语义相关专家(建模跨模态关系)和全局上下文专家(理解长程依赖),通过分层动态门控自适应整合,实现上下文感知的特征选择与模态加权。增强后的多模态表示通过伪标记注入与提示条件化无缝融入语言模型,使单一生成框架可处理分类与回归任务。采用参数高效的LoRA微调以保持计算效率。在双语基准数据集MELD、CHERMA、MOSEI、SIMS-V2上的大量实验表明,EGMF在准确率、跨语言鲁棒性及发现多模态情绪表达通用模式方面优于现有方法。

原文摘要 · Abstract (English)

Multimodal emotion understanding requires the integration of heterogeneous data sources, including text, audio, and visual modalities, while simultaneously addressing discrete emotion recognition and continuous sentiment analysis. We propose EGMF, a unified framework that combines expert-guided multimodal fusion with large language models to achieve superior performance across both tasks. At the core of our framework is a multi-scale expert network, comprising a local expert for capturing subtle emotional nuances, a semantic correlation expert for modeling cross-modal relationships, and a global context expert for understanding long-range dependencies. These experts are adaptively integrated via hierarchical dynamic gating, enabling context-aware feature selection and modality weighting. The enhanced multimodal representations are seamlessly incorporated into the language model through pseudo token injection and prompt-based conditioning, allowing a single generative framework to handle both classification and regression tasks. We employ parameter-efficient LoRA fine-tuning to maintain computational efficiency. Extensive experiments on bilingual benchmark datasets (MELD, CHERMA, MOSEI, SIMS-V2) demonstrate that EGMF outperforms state-of-the-art methods in terms of accuracy, cross-lingual robustness, and the discovery of universal patterns in multimodal emotional expressions.

多模态情绪识别语言模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。