跨语料库验证让乌尔都语语音情感识别更真实可信
Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
- 用eGeMAPS和ComParE特征,结合逻辑回归与多层感知机分类
- 跨语料库评估中准确率比自语料库低最多13%,反映模型泛化能力
- 为低资源语言情感计算提供可靠评估范式,适合相关研究者参考
语音情感识别(SER)是实现情感智能人工智能的关键技术。尽管普遍具有挑战性,对乌尔都语等低资源语言而言尤为困难。本研究在跨语料库设置下探索乌尔都语语音情感识别,这一领域此前几乎未被研究。我们采用跨语料库评估框架,在三个不同的乌尔都语情感语音数据集上测试模型泛化能力。使用两种基于领域知识的声学特征集——eGeMAPS和ComParE,将语音信号转换为特征向量,并输入逻辑回归与多层感知机分类器。通过未加权平均召回率(UAR)评估分类性能,同时考虑类别不平衡问题。结果显示,自语料库验证常高估性能,跨语料库评估结果最高可低13%,表明跨语料库评估更能真实反映模型鲁棒性。本研究强调了跨语料库验证在乌尔都语语音情感识别中的重要性,其成果有助于推动代表性不足语言群体的情感计算研究。
原文摘要 · Abstract (English)
Speech Emotion Recognition (SER) is a key affective computing technology that enables emotionally intelligent artificial intelligence. While SER is challenging in general, it is particularly difficult for low-resource languages such as Urdu. This study investigates Urdu SER in a cross-corpus setting, an area that has remained largely unexplored. We employ a cross-corpus evaluation framework across three different Urdu emotional speech datasets to test model generalization. Two standard domain-knowledge based acoustic feature sets, eGeMAPS and ComParE, are used to represent speech signals as feature vectors which are then passed to Logistic Regression and Multilayer Perceptron classifiers. Classification performance is assessed using unweighted average recall (UAR) whilst considering class-label imbalance. Results show that Self-corpus validation often overestimates performance, with UAR exceeding cross-corpus evaluation by up to 13%, underscoring that cross-corpus evaluation offers a more realistic measure of model robustness. Overall, this work emphasizes the importance of cross-corpus validation for Urdu SER and its implications contribute to advancing affective computing research for underrepresented language communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。