arXiv:2508.11371cs.SDeess.AS2025-08

用音频特征优化模型,拿下语音情感识别竞赛第一名

Speech Emotion Recognition Using Fine-Tuned DWFormer:A Study on Track 1 of the IERPChallenge 2024

  • 基于音频特征微调预训练模型DWFormer,融合数据增强与评分融合策略
  • 在IERP Challenge 2024 Track 1中取得第一名,性能领先参赛团队
  • 适合关注语音情感识别、模型微调与多模态融合的研究者

人工智能领域对情感识别高度关注。现有情感识别模型多聚焦于离散情绪标签的预测精度提升。鉴于人格特质与情感之间的直接关联,以及个体间主观情感表达的显著差异,IERP Challenge 2024 将人格特质纳入情感识别研究。本文报告了Fosafer团队在该挑战赛Track 1中的参赛方案。该任务主要关注音频中的情感识别,并提供文本与音频特征。在Track 1中,我们仅使用音频特征,通过数据增强与评分融合策略对预训练语音情感识别模型DWFormer进行微调,最终在所有参赛团队中排名第一。

原文摘要 · Abstract (English)

The field of artificial intelligence has a strong interest in the topic of emotion recognition. The majority of extant emotion recognition models are oriented towards enhancing the precision of discrete emotion label prediction. Given the direct relationship between human personality and emotion, as well as the significant inter-individual differences in subjective emotional expression, the IERP Challenge 2024 incorporates personality traits into emotion recognition research. This paper presents the Fosafer submissions to the Track 1 of the IERP Challenge 2024. This task primarily concerns the recognition of emotions in audio, while also providing text and audio features. In Track 1, we utilized exclusively audio-based features and fine-tuned a pre-trained speech emotion recognition model, DWFormer, through the integration of data augmentation and score fusion strategies, thereby achieving the first place among the participating teams.

语音情感识别模型微调音频特征竞赛第一

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。