构建首个丹麦语情感语音数据集,支持跨语言情感识别研究
EmoTale: An Enacted Speech-emotion Dataset in Danish
- 采集丹麦语与英语情感语音,标注表演性情绪标签
- 基于自监督模型嵌入,实现64.1%的无加权平均召回率
- 填补小语种情感语音数据空白,适合语音情感识别研究者
尽管已有多种常用语言的情感语音语料库,但小语种(如丹麦语)仍缺乏功能性数据集。据我们所知,1997年发布的丹麦情感语音(DES)是唯一一个丹麦语情感语音数据库。本文提出EmoTale,一个包含丹麦语和英语语音录音及其对应表演性情绪标注的语料库。通过使用自监督语音模型(SSLM)嵌入和openSMILE特征提取器,构建了针对EmoTale和参考数据集的语音情感识别(SER)模型。结果表明,嵌入特征优于手工设计特征。最佳模型在留一说话人交叉验证下,在EmoTale语料库上达到64.1%的无加权平均召回率(UAR),性能与DES相当。
原文摘要 · Abstract (English)
While multiple emotional speech corpora exist for commonly spoken languages, there is a lack of functional datasets for smaller (spoken) languages, such as Danish. To our knowledge, Danish Emotional Speech (DES), published in 1997, is the only other database of Danish emotional speech. We present EmoTale; a corpus comprising Danish and English speech recordings with their associated enacted emotion annotations. We demonstrate the validity of the dataset by investigating and presenting its predictive power using speech emotion recognition (SER) models. We develop SER models for EmoTale and the reference datasets using self-supervised speech model (SSLM) embeddings and the openSMILE feature extractor. We find the embeddings superior to the hand-crafted features. The best model achieves an unweighted average recall (UAR) of 64.1% on the EmoTale corpus using leave-one-speaker-out cross-validation, comparable to the performance on DES.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。