arXiv:2506.12325cs.SDcs.CL2025-06被引 1

通过图谱扩散重建缺失模态,提升对话情感识别准确率

GSDNet: Revisiting Incomplete Multimodal-Diffusion from Graph Spectrum Perspective for Conversation Emotion Recognition

  • 将噪声映射到图谱空间,仅调整特征值恢复缺失模态
  • 在多种模态缺失情况下达到当前最优情感识别效果
  • 适合处理多模态数据不全的现实场景,如语音或视频丢失

对话中的多模态情感识别(MERC)旨在通过视频、音频和文本等多种来源的语句信息推断说话人的情绪状态。相比单模态方法,融合不同模态的互补语义信息可获得更鲁棒的语句表征。然而,模态缺失问题严重限制了实际应用中的性能表现。近期研究分别利用图神经网络和扩散模型实现了较好的模态补全效果。受此启发,本文提出一种新型图谱扩散网络(GSDNet),将高斯噪声映射至缺失模态的图谱空间,并根据原始分布恢复缺失数据。与以往图扩散方法不同,GSDNet仅改变邻接矩阵的特征值,而非直接破坏邻接关系,从而在扩散过程中保留全局拓扑结构与关键谱特征。大量实验表明,GSDNet在多种模态缺失条件下均实现当前最优的情感识别性能。

原文摘要 · Abstract (English)

Multimodal emotion recognition in conversations (MERC) aims to infer the speaker's emotional state by analyzing utterance information from multiple sources (i.e., video, audio, and text). Compared with unimodality, a more robust utterance representation can be obtained by fusing complementary semantic information from different modalities. However, the modality missing problem severely limits the performance of MERC in practical scenarios. Recent work has achieved impressive performance on modality completion using graph neural networks and diffusion models, respectively. This inspires us to combine these two dimensions through the graph diffusion model to obtain more powerful modal recovery capabilities. Unfortunately, existing graph diffusion models may destroy the connectivity and local structure of the graph by directly adding Gaussian noise to the adjacency matrix, resulting in the generated graph data being unable to retain the semantic and topological information of the original graph. To this end, we propose a novel Graph Spectral Diffusion Network (GSDNet), which maps Gaussian noise to the graph spectral space of missing modalities and recovers the missing data according to its original distribution. Compared with previous graph diffusion methods, GSDNet only affects the eigenvalues of the adjacency matrix instead of destroying the adjacency matrix directly, which can maintain the global topological information and important spectral features during the diffusion process. Extensive experiments have demonstrated that GSDNet achieves state-of-the-art emotion recognition performance in various modality loss scenarios.

多模态图神经网络情感识别扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。