arXiv:2603.22345cs.AI2026-03被引 1

动态融合机制提升对话情绪识别,让模型针对不同情绪自适应调整参数。

Dynamic Fusion-Aware Graph Convolutional Neural Network for Multimodal Emotion Recognition in Conversations

  • 引入微分方程建模对话中情绪依赖的动态变化。
  • 通过全局信息向量生成提示,动态调节多模态特征融合方式。
  • 在两个公开数据集上表现更优,尤其适合复杂情绪分类任务。

对话中的多模态情绪识别(MERC)旨在从文本、音频等多源信息中理解说话人的情绪。现有方法虽利用图卷积网络(GCN)建模说话人间依赖关系,但通常对所有情绪类型使用固定参数处理多模态特征,忽略了模态融合的动态性,导致模型难以在特定情绪上表现突出。为此,本文提出动态融合感知图卷积神经网络(DF-GCN),将常微分方程引入GCN以捕捉话语交互网络中情绪依赖的动态特性,并利用话语全局信息向量(GIV)生成提示,引导多模态特征的动态融合。该设计使模型在推理时可针对不同情绪类别动态调整参数,实现更灵活的情绪分类与更强的泛化能力。在两个公开多模态对话数据集上的全面实验表明,所提DF-GCN显著优于基线模型,性能提升主要归因于动态融合机制。

原文摘要 · Abstract (English)

Multimodal emotion recognition in conversations (MERC) aims to identify and understand the emotions expressed by speakers during utterance interaction from multiple modalities (e.g., text, audio, images, etc.). Existing studies have shown that GCN can improve the performance of MERC by modeling dependencies between speakers. However, existing methods usually use fixed parameters to process multimodal features for different emotion types, ignoring the dynamics of fusion between different modalities, which forces the model to balance performance between multiple emotion categories, thus limiting the model's performance on some specific emotions. To this end, we propose a dynamic fusion-aware graph convolutional neural network (DF-GCN) for robust recognition of multimodal emotion features in conversations. Specifically, DF-GCN integrates ordinary differential equations into graph convolutional networks (GCNs) to {capture} the dynamic nature of emotional dependencies within utterance interaction networks and leverages the prompts generated by the global information vector (GIV) of the utterance to guide the dynamic fusion of multimodal features. This allows our model to dynamically change parameters when processing each utterance feature, so that different network parameters can be equipped for different emotion categories in the inference stage, thereby achieving more flexible emotion classification and enhancing the generalization ability of the model. Comprehensive experiments conducted on two public multimodal conversational datasets {confirm} that the proposed DF-GCN model delivers superior performance, benefiting significantly from the dynamic fusion mechanism introduced.

情绪识别图神经网络多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。