融合语音、文本与多模态模型,提升真实场景下情感识别准确率
ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge
- 结合语音、文本及语音-文本联合建模,利用大语言模型捕捉情感线索
- 针对类别不平衡问题,采用定制损失函数与多数投票决策机制
- 在166个参赛系统中位列第4,完整训练后达到当前最佳性能
自然语境下的语音情感识别(SER)仍面临内在变异性、多样录音条件和类别不平衡等挑战。作为聚焦这些复杂性的Interspeech自然语境SER挑战赛的参与者,我们提出Abhinaya系统,集成语音、文本及语音-文本模型。该方法对自监督语音大语言模型(SLLM)进行微调以提取语音表征,利用大语言模型(LLM)获取文本上下文,并通过结合SLLM的语音-文本建模捕捉细微情感信号。为缓解类别不平衡,采用定制化损失函数并基于多数投票生成分类决策。尽管其中一个模型未完全训练,系统仍于166个提交方案中排名第四;完整训练后,其性能超越已有公开结果,验证了该方法在真实场景下语音情感识别的有效性。
原文摘要 · Abstract (English)
Speech emotion recognition (SER) in naturalistic settings remains a challenge due to the intrinsic variability, diverse recording conditions, and class imbalance. As participants in the Interspeech Naturalistic SER Challenge which focused on these complexities, we present Abhinaya, a system integrating speech-based, text-based, and speech-text models. Our approach fine-tunes self-supervised and speech large language models (SLLM) for speech representations, leverages large language models (LLM) for textual context, and employs speech-text modeling with an SLLM to capture nuanced emotional cues. To combat class imbalance, we apply tailored loss functions and generate categorical decisions through majority voting. Despite one model not being fully trained, the Abhinaya system ranked 4th among 166 submissions. Upon completion of training, it achieved state-of-the-art performance among published results, demonstrating the effectiveness of our approach for SER in real-world conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。