用熵值筛选融合语音与文本情绪预测,提升识别准确率
Speech Emotion Recognition via Entropy-Aware Score Selection
- 基于声学与文本双流模型,通过熵阈值筛选高可信度预测
- 在IEMOCAP和MSP-IMPROV上准确率优于单模态系统
- 适合需要多模态情绪分析的语音交互场景
本文提出一种多模态语音情绪识别框架,采用熵感知得分选择策略融合语音与文本预测。主路径使用基于wav2vec2.0的声学模型,辅路径采用基于RoBERTa-XLM的情感分析模型,转录文本由Whisper-large-v3生成。提出基于熵与方差熵阈值的晚期得分融合方法,以克服主路径预测的置信度局限。通过情感映射策略将三种情感类别转换为四类目标情绪,实现多模态预测的一致融合。在IEMOCAP和MSP-IMPROV数据集上的实验表明,该方法相较于传统单模态系统具有显著且可靠的性能提升。
原文摘要 · Abstract (English)
In this paper, we propose a multimodal framework for speech emotion recognition that leverages entropy-aware score selection to combine speech and textual predictions. The proposed method integrates a primary pipeline that consists of an acoustic model based on wav2vec2.0 and a secondary pipeline that consists of a sentiment analysis model using RoBERTa-XLM, with transcriptions generated via Whisper-large-v3. We propose a late score fusion approach based on entropy and varentropy thresholds to overcome the confidence constraints of primary pipeline predictions. A sentiment mapping strategy translates three sentiment categories into four target emotion classes, enabling coherent integration of multimodal predictions. The results on the IEMOCAP and MSP-IMPROV datasets show that the proposed method offers a practical and reliable enhancement over traditional single-modality systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。