用声带和口腔运动数据提升语音情绪识别准确率
Speech Emotion Recognition with Phonation Excitation Information and Articulatory Kinematics
- 融合声带振动与口腔动作生理数据,增强情绪识别能力
- 在自建数据集STEM-E2VA上,生理信息使识别率显著提升
- 可通过语音反演获得生理数据,适合实际场景应用
语音情绪识别(SER)得益于深度学习方法已取得显著进展,文本信息进一步提升了性能。然而,现有研究较少关注语音生成过程中的生理信息,而这些信息也蕴含说话人特征,包括情绪状态。为弥补这一空白,我们开展实验,探究声带激励信息(通过电声门图,EGG获取)与发音运动学信息(通过电磁发音仪,EMA获取)在SER中的潜力。由于此类数据稀缺,我们构建了包含音频与生理信号的标注情感数据集STEM-E2VA。此外,我们还利用从语音中反演得到的生理数据进行情绪识别,以验证其在真实场景中的可行性。实验结果表明,引入语音生成的生理信息能有效提升SER性能,具备实际应用前景。
原文摘要 · Abstract (English)
Speech emotion recognition (SER) has advanced significantly for the sake of deep-learning methods, while textual information further enhances its performance. However, few studies have focused on the physiological information during speech production, which also encompasses speaker traits, including emotional states. To bridge this gap, we conducted a series of experiments to investigate the potential of the phonation excitation information and articulatory kinematics for SER. Due to the scarcity of training data for this purpose, we introduce a portrayed emotional dataset, STEM-E2VA, which includes audio and physiological data such as electroglottography (EGG) and electromagnetic articulography (EMA). EGG and EMA provide information of phonation excitation and articulatory kinematics, respectively. Additionally, we performed emotion recognition using estimated physiological data derived through inversion methods from speech, instead of collected EGG and EMA, to explore the feasibility of applying such physiological information in real-world SER. Experimental results confirm the effectiveness of incorporating physiological information about speech production for SER and demonstrate its potential for practical use in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。