arXiv:2501.15063cs.CL2025-01被引 9

融合跨模态上下文与自适应图卷积,提升对话情绪识别准确率

Cross-modal Context Fusion and Adaptive Graph Convolutional Network for Multimodal Conversational Emotion Recognition

论文配图:Cross-modal Context Fusion and Adaptive Graph Convolutional Network for Multimodal Conversational Emotion Recognition
图 1 · 摘自论文原文
  • 通过跨模态对齐与上下文融合减少多模态干扰
  • 构建对话关系图捕捉说话人间依赖与自我依赖
  • 适用于人机交互、医疗等需要精准情绪理解的场景

情绪识别在人机交互、营销、医疗等领域具有广泛应用。近年来深度学习技术为情绪识别提供了新方法。尽管已有多种多模态情绪识别方法提出,但这些方法往往忽视不同输入模态间的相互干扰,且较少关注说话人之间的方向性对话关系。为此,本文提出一种新型多模态情绪识别方法,包含跨模态上下文融合模块、自适应图卷积编码模块和情绪分类模块。跨模态上下文模块包含跨模态对齐与上下文融合单元,用于降低多模态间相互干扰带来的噪声。自适应图卷积模块构建对话关系图,以提取说话人之间的依赖关系与自我依赖关系。所提模型在公开基准数据集上超越多项先进方法,实现了高情绪识别准确率。

原文摘要 · Abstract (English)

Emotion recognition has a wide range of applications in human-computer interaction, marketing, healthcare, and other fields. In recent years, the development of deep learning technology has provided new methods for emotion recognition. Prior to this, many emotion recognition methods have been proposed, including multimodal emotion recognition methods, but these methods ignore the mutual interference between different input modalities and pay little attention to the directional dialogue between speakers. Therefore, this article proposes a new multimodal emotion recognition method, including a cross modal context fusion module, an adaptive graph convolutional encoding module, and an emotion classification module. The cross modal context module includes a cross modal alignment module and a context fusion module, which are used to reduce the noise introduced by mutual interference between different input modalities. The adaptive graph convolution module constructs a dialogue relationship graph for extracting dependencies and self dependencies between speakers. Our model has surpassed some state-of-the-art methods on publicly available benchmark datasets and achieved high recognition accuracy.

情绪识别多模态图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。