用跨注意力融合多特征,提升跨语言语音情感识别准确率
Leveraging Cross-Attention Transformer and Multi-Feature Fusion for Cross-Linguistic Speech Emotion Recognition
- 融合HuBERT、MFCC和韵律特征,通过跨注意力机制提取联合表示
- 在7个跨语言数据集上平均准确率达78.75%,德国语EMODB达88.69%
- 小样本微调即可适配新语言,适合多语言交互系统研发
语音情感识别(SER)对提升人机交互体验至关重要。跨语言语音情感识别(CLSER)因语言间声学与语言特征差异大而极具挑战。本文提出HuMP-CAT方法,融合HuBERT、MFCC及韵律特征,并通过交叉注意力变换器(CAT)在特征提取阶段实现多模态融合。利用迁移学习,以IEMOCAP为源数据集训练源模型,在7个涵盖英语、德语、西班牙语、意大利语和中文的跨语言数据集上进行评估。仅用少量目标语言语音微调,该方法在7个数据集上平均准确率达78.75%,其中德语EMODB达到88.69%,意大利语EMOVO达79.48%。实验证明,HuMP-CAT在多语言场景下优于现有方法。
原文摘要 · Abstract (English)
Speech Emotion Recognition (SER) plays a crucial role in enhancing human-computer interaction. Cross-Linguistic SER (CLSER) has been a challenging research problem due to significant variability in linguistic and acoustic features of different languages. In this study, we propose a novel approach HuMP-CAT, which combines HuBERT, MFCC, and prosodic characteristics. These features are fused using a cross-attention transformer (CAT) mechanism during feature extraction. Transfer learning is applied to gain from a source emotional speech dataset to the target corpus for emotion recognition. We use IEMOCAP as the source dataset to train the source model and evaluate the proposed method on seven datasets in five languages (e.g., English, German, Spanish, Italian, and Chinese). We show that, by fine-tuning the source model with a small portion of speech from the target datasets, HuMP-CAT achieves an average accuracy of 78.75% across the seven datasets, with notable performance of 88.69% on EMODB (German language) and 79.48% on EMOVO (Italian language). Our extensive evaluation demonstrates that HuMP-CAT outperforms existing methods across multiple target languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。