arXiv:2601.18010eess.AScs.SD2026-01中稿 · ICASSP 2026被引 3

同时处理情绪标注与模态间的不确定性,提升语音文本情感识别准确率。

AmbER$^2$: Dual Ambiguity-Aware Emotion Recognition Applied to Speech and Text

  • 采用师生架构与分布训练目标,双路建模评分者与模态间模糊性。
  • 在IEMOCAP上各项指标提升3.8%~20.3%,尤其对高不确定样本效果显著。
  • 适合构建鲁棒情感识别系统的研究者,尤其是多模态场景应用者。

情绪识别本身具有固有模糊性,源于评分者之间意见不一及语音与文本等模态间的差异。现有研究多关注通过标签分布建模评分者模糊性,但模态模糊性仍被忽视,多数多模态方法仅依赖简单特征融合,未显式处理模态冲突。本文提出AmbER$^2$,一种双模糊性感知框架,通过师生架构与分布级训练目标,同时建模评分者层级与模态层级的模糊性。在IEMOCAP与MSP-Podcast数据集上的评估表明,AmbER$^2$在分布保真度上持续优于传统交叉熵基线,性能可媲美或超越近期先进系统。例如,在IEMOCAP上,巴塔查里亚系数相对提升20.3%(0.83 vs. 0.69),R²提升13.6%(0.67 vs. 0.59),准确率提升3.8%(0.683 vs. 0.658),F1提升4.5%(0.675 vs. 0.646)。进一步分析显示,显式建模模糊性在高不确定样本中尤为有效。结果凸显了联合处理评分者与模态模糊性对构建鲁棒情绪识别系统的重要性。

原文摘要 · Abstract (English)

Emotion recognition is inherently ambiguous, with uncertainty arising both from rater disagreement and from discrepancies across modalities such as speech and text. There is growing interest in modeling rater ambiguity using label distributions. However, modality ambiguity remains underexplored, and multimodal approaches often rely on simple feature fusion without explicitly addressing conflicts between modalities. In this work, we propose AmbER$^2$, a dual ambiguity-aware framework that simultaneously models rater-level and modality-level ambiguity through a teacher-student architecture with a distribution-wise training objective. Evaluations on IEMOCAP and MSP-Podcast show that AmbER$^2$ consistently improves distributional fidelity over conventional cross-entropy baselines and achieves performance competitive with, or superior to, recent state-of-the-art systems. For example, on IEMOCAP, AmbER$^2$ achieves relative improvements of 20.3% on Bhattacharyya coefficient (0.83 vs. 0.69), 13.6% on R$^2$ (0.67 vs. 0.59), 3.8% on accuracy (0.683 vs. 0.658), and 4.5% on F1 (0.675 vs. 0.646). Further analysis across ambiguity levels shows that explicitly modeling ambiguity is particularly beneficial for highly uncertain samples. These findings highlight the importance of jointly addressing rater and modality ambiguity when building robust emotion recognition systems.

情感识别多模态模糊建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。