针对自然语境下的情绪识别挑战,提出简洁高效的SAILER系统。
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
- 基于语音基础模型与优化学习目标,设计可复现的情绪识别框架。
- 单模型宏平均F1超0.4,超越95%参赛作品,集成后进入前三。
- 适合关注语音情感分析、数据增强与工程优化的研究者参考。
语音情绪识别(SER)在自然表达情绪场景下仍具挑战性,主要难点在于情绪标注的主观性及数据集中情绪标签分布不均。本文介绍了为参与INTERSPEECH 2025情感识别挑战赛(任务1)所开发的SAILER系统。该挑战数据集包含来自播客的真实情感语音,是研究不平衡与主观标注的重要资源。系统设计注重简洁性、可复现性与有效性,重点优化了建模方法、学习目标、数据增强及工程实现。结果表明,即使不使用集成,单一模型的宏平均F1得分超过0.4,超越95%以上提交结果;三模型集成进一步提升性能,取得前三名成绩。代码已开源:https://github.com/tiantiaf0627/vox-profile-release。
原文摘要 · Abstract (English)
Speech emotion recognition (SER), particularly for naturally expressed emotions, remains a challenging computational task. Key challenges include the inherent subjectivity in emotion annotation and the imbalanced distribution of emotion labels in datasets. This paper introduces the \texttt{SAILER} system developed for participation in the INTERSPEECH 2025 Emotion Recognition Challenge (Task 1). The challenge dataset, which contains natural emotional speech from podcasts, serves as a valuable resource for studying imbalanced and subjective emotion annotations. Our system is designed to be simple, reproducible, and effective, highlighting critical choices in modeling, learning objectives, data augmentation, and engineering choices. Results show that even a single system (without ensembling) can outperform more than 95\% of the submissions, with a Macro-F1 score exceeding 0.4. Moreover, an ensemble of three systems further improves performance, achieving a competitively ranked score (top-3 performing team). Our model is at: https://github.com/tiantiaf0627/vox-profile-release.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。