arXiv:2507.05624cs.AI2025-07被引 2

用注意力扩散模型补全缺失模态特征,提升情感意图识别效果

ADMC: Attention-based Diffusion Model for Missing Modalities Feature Completion

  • 分模态独立训练,避免特征耦合
  • 在IEMOCAP和MIntRec上达顶尖性能
  • 支持缺模态与全模态场景,通用性强

多模态情感与意图识别对人机交互至关重要,旨在分析用户的语音、文本和视觉信息以预测情绪或意图。一个关键挑战是传感器故障或数据不完整导致的模态缺失。传统重建方法常因过度耦合和生成不精准而表现不佳。为此,我们提出注意力扩散模型用于缺失模态特征补全(ADMC)。该框架为各模态独立训练特征提取网络,保留其独特性并避免过耦合。注意力扩散网络(ADN)生成的缺失模态特征能紧密贴合真实多模态分布,在各类缺失场景下均显著提升识别性能。此外,ADN的跨模态生成能力在全模态情况下也带来性能增益。在IEMOCAP和MIntRec基准上取得当前最优结果,验证了其在缺失与完整模态场景下的有效性。

原文摘要 · Abstract (English)

Multimodal emotion and intent recognition is essential for automated human-computer interaction, It aims to analyze users' speech, text, and visual information to predict their emotions or intent. One of the significant challenges is that missing modalities due to sensor malfunctions or incomplete data. Traditional methods that attempt to reconstruct missing information often suffer from over-coupling and imprecise generation processes, leading to suboptimal outcomes. To address these issues, we introduce an Attention-based Diffusion model for Missing Modalities feature Completion (ADMC). Our framework independently trains feature extraction networks for each modality, preserving their unique characteristics and avoiding over-coupling. The Attention-based Diffusion Network (ADN) generates missing modality features that closely align with authentic multimodal distribution, enhancing performance across all missing-modality scenarios. Moreover, ADN's cross-modal generation offers improved recognition even in full-modality contexts. Our approach achieves state-of-the-art results on the IEMOCAP and MIntRec benchmarks, demonstrating its effectiveness in both missing and complete modality scenarios.

多模态特征补全扩散模型情感识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。