提出混合扩散模型修复缺失多模态情绪数据,提升识别鲁棒性。
RoHyDR: Robust Hybrid Diffusion Recovery for Incomplete Multimodal Emotion Recognition
- 用扩散生成器从噪声中恢复缺失模态的特征表示。
- 结合对抗学习重建融合特征与语义内容,支持多层级修复。
- 在两个基准上优于现有方法,适合真实场景下的情绪识别任务。
多模态情绪识别通过融合多源数据分析情绪,但现实中的噪声或传感器故障常导致数据缺失或损坏,形成不完整多模态情绪识别(IMER)挑战。本文提出鲁棒混合扩散恢复框架RoHyDR,实现缺失模态在单模态、多模态、特征和语义层面的恢复。针对缺失模态的单模态表征恢复,RoHyDR利用基于扩散的生成器,从高斯噪声生成分布一致且语义对齐的表征,以可用模态作为条件。对于多模态融合恢复,引入对抗学习生成逼真的融合表征并恢复缺失语义内容。进一步设计多阶段优化策略,提升训练稳定性和效率。相比以往方法,RoHyDR的混合扩散与对抗学习机制可同时在特征与语义层实现单模态与多模态融合的缺失信息恢复,有效缓解次优优化带来的性能下降。在两个常用多模态情绪识别基准上的全面实验表明,所提方法优于当前最优的IMER方法,在多种缺失模态场景下均表现稳健。代码将在录用后公开。
原文摘要 · Abstract (English)
Multimodal emotion recognition analyzes emotions by combining data from multiple sources. However, real-world noise or sensor failures often cause missing or corrupted data, creating the Incomplete Multimodal Emotion Recognition (IMER) challenge. In this paper, we propose Robust Hybrid Diffusion Recovery (RoHyDR), a novel framework that performs missing-modality recovery at unimodal, multimodal, feature, and semantic levels. For unimodal representation recovery of missing modalities, RoHyDR exploits a diffusion-based generator to generate distribution-consistent and semantically aligned representations from Gaussian noise, using available modalities as conditioning. For multimodal fusion recovery, we introduce adversarial learning to produce a realistic fused multimodal representation and recover missing semantic content. We further propose a multi-stage optimization strategy that enhances training stability and efficiency. In contrast to previous work, the hybrid diffusion and adversarial learning-based recovery mechanism in RoHyDR allows recovery of missing information in both unimodal representation and multimodal fusion, at both feature and semantic levels, effectively mitigating performance degradation caused by suboptimal optimization. Comprehensive experiments conducted on two widely used multimodal emotion recognition benchmarks demonstrate that our proposed method outperforms state-of-the-art IMER methods, achieving robust recognition performance under various missing-modality scenarios. Our code will be made publicly available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。