arXiv:2506.15754cs.SDeess.AS2025-06被引 5

用注意力池化提升语音情感识别准确率并揭示情绪关键帧

Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization

  • 设计多查询多头注意力池化,比平均池化提升3.5个百分点
  • 仅15%的音频帧包含80%的情绪线索,情绪信息高度集中
  • 发现非语言发声和夸张发音被优先关注,贴近人类感知机制

当前语音情感识别(SER)的先进变换器模型依赖时间特征聚合,但高级池化方法仍缺乏系统探索。我们系统性地评估了多种池化策略,其中多查询多头注意力统计池化相比平均池化在宏观F1上提升3.5个百分点。注意力分析显示,仅有15%的音频帧承载80%的情绪线索,揭示情绪信息具有高度局部化特征。对高注意力帧的分析表明,非语言声学特征与超清晰发音在池化过程中被显著优先处理,符合人类感知策略。研究结果表明,注意力池化不仅提升了SER性能,还提供了可解释的情绪定位机制。在Interspeech 2025自然情境语音情感识别挑战赛中,该方法取得0.3649的宏观F1得分。

原文摘要 · Abstract (English)

State-of-the-art transformer models for Speech Emotion Recognition (SER) rely on temporal feature aggregation, yet advanced pooling methods remain underexplored. We systematically benchmark pooling strategies, including Multi-Query Multi-Head Attentive Statistics Pooling, which achieves a 3.5 percentage point macro F1 gain over average pooling. Attention analysis shows 15 percent of frames capture 80 percent of emotion cues, revealing a localized pattern of emotional information. Analysis of high-attention frames reveals that non-linguistic vocalizations and hyperarticulated phonemes are disproportionately prioritized during pooling, mirroring human perceptual strategies. Our findings position attentive pooling as both a performant SER mechanism and a biologically plausible tool for explainable emotion localization. On Interspeech 2025 Speech Emotion Recognition in Naturalistic Conditions Challenge, our approach obtained a macro F1 score of 0.3649.

语音情感识别注意力机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。