arXiv:2502.10435cs.CVcs.AI2025-02IJCAI被引 4

解决多人多模态情绪识别中数据缺失问题,提升模型鲁棒性。

RAMer: Reconstruction-based Adversarial Model for Multi-party Multi-modal Multi-label Emotion Recognition

  • 基于重建对抗机制,融合模态共性和特异性特征
  • 在三个基准上达到当前最优性能,尤其在不完整数据下表现优异
  • 适合处理真实场景中缺失语音/文本的多人情绪识别任务

传统多模态多标签情绪识别(MMER)假设视觉、文本和语音模态均完整可用。但在真实多人场景中,非发言者常缺失语音和文本信息,导致模型性能显著下降。现有方法通常将异构模态统一为单一表示,忽视各模态独特性。为此,我们提出RAMer(基于重建的对抗式情绪识别模型),通过探索模态共性与特异性,并利用对比学习增强的重构特征,缓解数据不完整问题,提升特征质量。同时引入个性辅助任务,借助模态级注意力补全缺失模态,改善情绪推理。为进一步捕捉标签与模态间依赖关系,提出堆叠打乱策略,强化标签与模态特定特征的关联。在MEmoR、CMU-MOSEI和$M^3ED$三个基准上的实验表明,RAMer在双人及多人MMER场景中均达到先进水平。

原文摘要 · Abstract (English)

Conventional Multi-modal multi-label emotion recognition (MMER) assumes complete access to visual, textual, and acoustic modalities. However, real-world multi-party settings often violate this assumption, as non-speakers frequently lack acoustic and textual inputs, leading to a significant degradation in model performance. Existing approaches also tend to unify heterogeneous modalities into a single representation, overlooking each modality's unique characteristics. To address these challenges, we propose RAMer (Reconstruction-based Adversarial Model for Emotion Recognition), which refines multi-modal representations by not only exploring modality commonality and specificity but crucially by leveraging reconstructed features, enhanced by contrastive learning, to overcome data incompleteness and enrich feature quality. RAMer also introduces a personality auxiliary task to complement missing modalities using modality-level attention, improving emotion reasoning. To further strengthen the model's ability to capture label and modality interdependency, we propose a stack shuffle strategy to enrich correlations between labels and modality-specific features. Experiments on three benchmarks, i.e., MEmoR, CMU-MOSEI, and $M^3ED$, demonstrate that RAMer achieves state-of-the-art performance in dyadic and multi-party MMER scenarios.

情绪识别多模态数据缺失对抗学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。