arXiv:2409.05566eess.AS2024-09中稿 · publication at IEE…被引 5

用语音语义与声学特征联合建模,提升情绪识别效果。

Leveraging Content and Acoustic Representations for Speech Emotion Recognition

  • 双编码器设计,分别捕捉语音语义和声学特征
  • 仅用无标签语音训练,跨8个数据集表现最优
  • 轻量级分类器即可实现高效情绪识别,适合资源受限场景

语音情绪识别(SER)因难以从语音中提取情感表征而具有挑战性,且标注数据稀缺导致大模型易过拟合。本文提出CARE(情感的语义与声学表征),采用双编码方案,强调语音的语义与声学特性。语义编码器通过句级文本表征的蒸馏进行训练,声学编码器则预测语音信号的帧级低层特征。该方法基于仅在无监督原始语音上训练的基线尺寸模型,配合轻量级下游分类器,在多个数据集上表现出色。与多种自监督模型及基于大语言模型的方法对比,CARE在8个多样化数据集上的平均性能最佳。并通过消融实验验证了各设计选择的重要性。

原文摘要 · Abstract (English)

Speech emotion recognition (SER), the task of identifying the expression of emotion from spoken content, is challenging due to the difficulty in extracting representations that capture emotional attributes from speech. The scarcity of labeled datasets further complicates the challenge where large models are prone to over-fitting. In this paper, we propose CARE (Content and Acoustic Representations of Emotions), where we design a dual encoding scheme which emphasizes semantic and acoustic factors of speech. While the semantic encoder is trained using distillation from utterance-level text representations, the acoustic encoder is trained to predict low-level frame-wise features of the speech signal. The proposed dual encoding scheme is a base-sized model trained only on unsupervised raw speech. With a simple light-weight classification model trained on the downstream task, we show that the CARE embeddings provide effective emotion recognition on a variety of datasets. We compare the proposal with several other self-supervised models as well as recent large-language model based approaches. In these evaluations, the proposed CARE is shown to be the best performing model based on average performance across 8 diverse datasets. We also conduct several ablation studies to analyze the importance of various design choices.

语音识别情绪识别自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。