arXiv:2506.02088cs.SDcs.CL2025-06被引 3

用语音韵律与图注意力融合提升自然语境下情绪识别准确率

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

  • 结合音高量化与预训练音频标签模型增强语音特征
  • 在测试集上达39.79%宏平均F1分数,验证集42.20%
  • 适合关注真实场景情绪识别与多模态融合的研究者

在自然、自发的语音中训练情绪识别模型极具挑战,因情绪表达微妙且真实音频具有不可预测性。本文针对INTERSPEECH 2025自然语境语音情绪识别挑战,提出一种鲁棒系统,聚焦于类别化情绪识别。方法融合前沿音频模型与富含韵律和谱学线索的文本特征。特别地,研究了基频(F0)量化及预训练音频标记模型的使用效果。采用集成模型提升鲁棒性。在官方测试集上,系统取得39.79%的宏平均F1分数(验证集42.20%)。结果表明该方法有效,融合技术分析证实图注意力网络的优越性。源代码已公开。

原文摘要 · Abstract (English)

Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the INTERSPEECH 2025 Speech Emotion Recognition in Naturalistic Conditions Challenge, focusing on categorical emotion recognition. Our method combines state-of-the-art audio models with text features enriched by prosodic and spectral cues. In particular, we investigate the effectiveness of Fundamental Frequency (F0) quantization and the use of a pretrained audio tagging model. We also employ an ensemble model to improve robustness. On the official test set, our system achieved a Macro F1-score of 39.79% (42.20% on validation). Our results underscore the potential of these methods, and analysis of fusion techniques confirmed the effectiveness of Graph Attention Networks. Our source code is publicly available.

情绪识别语音分析图神经网络韵律特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。