解决语音合成中非语言发声的说话人验证难题,提升跨模态识别准确率。
Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach

- 融合自监督特征与专家混合模型,实现语音与非语言发声统一建模。
- 在10类非语言发声上将跨域错误率从38.93%降至22.66%。
- 适用于需高保真说话人一致性的语音合成与语音转换系统评测。
随着表达性文本转语音(TTS)和语音转换(VC)系统越来越多地生成非语言发声(NVVs)以增强自然度,可靠的身份验证(SV)对于客观评估语音与非语言段落间身份一致性变得至关重要。然而,现有SV系统对NVVs泛化能力差,且在NVV数据上微调会导致语音性能灾难性遗忘。我们首次系统研究了10种类型的非语言发声,并提出一种框架:结合冻结的Data2Vec自监督特征与ECAPA-TDNN,通过带有学习型领域感知路由的专家混合(MoE)模块进行增强。通过对语音输入施加预训练教师的条件蒸馏损失,保留语音到语音的准确性;同时使用对比损失弥合语音与非语言发声之间的域差距。相比预训练基线,该方法将语音到非语言发声的等错误率(EER)从38.93%降低至22.66%,并通过蒸馏使语音EER从13.17%提升至9.24%。
原文摘要 · Abstract (English)
As expressive text-to-speech (TTS) and voice conversion (VC) systems increasingly generate non-verbal vocalizations (NVVs) to enhance naturalness, reliable speaker verification (SV) becomes essential to objectively assess identity consistency across both verbal and non-verbal segments. Yet current SV systems generalize poorly to NVVs, and fine-tuning on NVV data causes catastrophic forgetting of speech performance. We present the first systematic study across 10 NVV types and propose a framework combining frozen Data2Vec self-supervised features with ECAPA-TDNN, enhanced by a Mixture of Experts (MoE) module with learned domain-aware routing. A conditional distillation loss on speech inputs via a pretrained teacher retains speech-to-speech accuracy, while a contrastive loss bridges the speech-NVV domain gap. Our method reduces speech-NVV EER from 38.93% to 22.66% over a pretrained baseline, and improves speech EER from 13.17% to 9.24% via distillation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。