用生成模型合成情感语音,提升语音识别在情绪变化下的准确率。
Improving speaker verification robustness with synthetic emotional utterances
- 用CycleGAN生成特定说话人的带情绪语音,保持音色不变。
- 训练时加入合成数据,使模型在情绪语音上误识率降低3.64%。
- 适合需要高鲁棒性语音验证的智能客服、安全认证场景。
语音验证(SV)系统用于确认一段语音是否来自特定说话人,广泛应用于个性化服务中。然而,现有模型在情绪化语音上的表现远差于中性语音,主要因标注的情绪语音数据稀缺,导致难以建模跨情绪的稳定说话人表征。为此,本文提出一种基于CycleGAN的数据增强方法,为每位说话人合成情感语音片段,同时保留其独特声纹特征。实验表明,使用该合成数据训练的模型,在情绪语音验证任务中显著优于基线模型,等错误率(EER)相对降低最多达3.64%。
原文摘要 · Abstract (English)
A speaker verification (SV) system offers an authentication service designed to confirm whether a given speech sample originates from a specific speaker. This technology has paved the way for various personalized applications that cater to individual preferences. A noteworthy challenge faced by SV systems is their ability to perform consistently across a range of emotional spectra. Most existing models exhibit high error rates when dealing with emotional utterances compared to neutral ones. Consequently, this phenomenon often leads to missing out on speech of interest. This issue primarily stems from the limited availability of labeled emotional speech data, impeding the development of robust speaker representations that encompass diverse emotional states. To address this concern, we propose a novel approach employing the CycleGAN framework to serve as a data augmentation method. This technique synthesizes emotional speech segments for each specific speaker while preserving the unique vocal identity. Our experimental findings underscore the effectiveness of incorporating synthetic emotional data into the training process. The models trained using this augmented dataset consistently outperform the baseline models on the task of verifying speakers in emotional speech scenarios, reducing equal error rate by as much as 3.64% relative.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。