用零样本语音合成增强短语音说话人验证,不需重训练就提升准确率。
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
- 在测试时用零样本语音合成生成新语音,扩充数据
- 短语音下等错误率降低10%-16%,效果显著
- 适合研究说话人验证与语音合成交叉方向的学者
短语音说话人验证因语音片段信息有限而面临挑战,影响准确性与可靠性。近年来,零样本文本到语音(ZS-TTS)系统在保留说话人身份方面取得显著进展。本研究首次探索将ZS-TTS系统用于测试时的数据增强。我们在VoxCeleb 1数据集上评估了三种先进的预训练ZS-TTS模型:NatureSpeech 3、CosyVoice和MaskGCT。实验结果表明,结合真实与合成语音样本可使所有语音时长下的等错误率(EER)相对降低10%-16%,尤其在短语音上表现突出,且无需重训练现有系统。然而分析发现,较长的合成语音未能像真实语音那样有效降低EER。这些结果揭示了使用ZS-TTS进行测试时说话人验证的潜力与挑战,为未来研究提供重要参考。
原文摘要 · Abstract (English)
Short-utterance speaker verification presents significant challenges due to the limited information in brief speech segments, which can undermine accuracy and reliability. Recently, zero-shot text-to-speech (ZS-TTS) systems have made considerable progress in preserving speaker identity. In this study, we explore, for the first time, the use of ZS-TTS systems for test-time data augmentation for speaker verification. We evaluate three state-of-the-art pre-trained ZS-TTS systems, NatureSpeech 3, CosyVoice, and MaskGCT, on the VoxCeleb 1 dataset. Our experimental results show that combining real and synthetic speech samples leads to 10%-16% relative equal error rate (EER) reductions across all durations, with particularly notable improvements for short utterances, all without retraining any existing systems. However, our analysis reveals that longer synthetic speech does not yield the same benefits as longer real speech in reducing EERs. These findings highlight the potential and challenges of using ZS-TTS for test-time speaker verification, offering insights for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。