对比多种预训练模型,发现专注语音副语言特征的模型更适合跨语料库情感识别。
Rethinking Cross-Corpus Speech Emotion Recognition Benchmarking: Are Paralinguistic Pre-Trained Representations Sufficient?
- 用多类预训练模型对比,重点评估语音副语言特征模型表现
- TRILLsson模型在跨语料库情感识别中准确率领先其他模型
- 研究提醒应重视副语言特征模型,提升评测可靠性
近期用于评估跨语料库语音情感识别(SER)的预训练模型(PTM)基准忽略了专为语音副语言处理(PSP)设计的模型,引发可信度质疑,因为情感识别本质上是副语言任务。本文假设专注于PSP的PTM在跨语料库SER中表现更优。为此,我们分析了包括副语言、单语、多语及说话人识别在内的多种SOTA PTM表示。结果证实,TRILLsson(一种副语言PTM)性能优于其他模型,强化了在跨语料库SER基准中纳入PSP聚焦模型的必要性。本研究提升了基准可信度,并为可靠跨语料库SER的PTM评估提供指导。
原文摘要 · Abstract (English)
Recent benchmarks evaluating pre-trained models (PTMs) for cross-corpus speech emotion recognition (SER) have overlooked PTM pre-trained for paralinguistic speech processing (PSP), raising concerns about their reliability, since SER is inherently a paralinguistic task. We hypothesize that PSP-focused PTM will perform better in cross-corpus SER settings. To test this, we analyze state-of-the-art PTMs representations including paralinguistic, monolingual, multilingual, and speaker recognition. Our results confirm that TRILLsson (a paralinguistic PTM) outperforms others, reinforcing the need to consider PSP-focused PTMs in cross-corpus SER benchmarks. This study enhances benchmark trustworthiness and guides PTMs evaluations for reliable cross-corpus SER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。