对比三种声纹嵌入在零样本多说话人语音合成中的效果
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
- 固定合成框架,测试H/ASP、x-vector和ECAPA-TDNN三种编码器
- 原生H/ASP编码器在主观与客观评价中均最优,ECAPA优于x-vector
- 提醒跨任务复用声纹模型需实证验证,适合语音合成研究者参考
零样本多说话人文本转语音系统依赖声纹嵌入,仅通过一段参考语音即可合成未见说话人的语音。尽管已有多种声纹嵌入用于说话人识别,其在零样本语音合成中的有效性仍缺乏充分研究。本文采用基于YourTTS的合成系统,在相同捷克读音数据集上训练,固定合成框架,对比原生H/ASP编码器、x-vector嵌入与ECAPA-TDNN嵌入的表现。在24个域外目标说话人上进行主客观评估:主观测试聚焦说话人相似度,客观评估则计算合成语音与真实语音提取的声纹嵌入间的余弦距离。结果表明,原生H/ASP编码器始终表现最佳,ECAPA-TDNN优于x-vector。该结果提示,尽管ECAPA-TDNN在说话人识别中广受青睐,但在本配置下并未提升零样本语音合成中的说话人相似度。研究强调了在语音合成中复用识别嵌入时进行实证评估的重要性,并为后续比较提供了可复现框架。
原文摘要 · Abstract (English)
Zero-shot multi-speaker text-to-speech (TTS) systems rely on speaker embeddings to synthesize speech in the voice of an unseen speaker, using only a short reference utterance. While many speaker embeddings have been developed for speaker recognition, their relative effectiveness in zero-shot TTS remains underexplored. In this work, we employ a YourTTS-based TTS system to compare three different speaker encoders - YourTTS's original H/ASP encoder, x-vector embeddings, and ECAPA-TDNN embeddings - within an otherwise fixed zero-shot TTS framework. All models were trained on the same dataset of Czech read speech and evaluated on 24 out-of-domain target speakers using both subjective and objective methods. The subjective evaluation was conducted via a listening test focused on speaker similarity, while the objective evaluation measured cosine distances between speaker embeddings extracted from synthesized and real utterances. Across both evaluations, the original H/ASP encoder consistently outperformed the alternatives, with ECAPA-TDNN showing better results than x-vectors. These findings suggest that, despite the popularity of ECAPA-TDNN in speaker recognition, it does not necessarily offer improvements for speaker similarity in zero-shot TTS in this configuration. Our study highlights the importance of empirical evaluation when reusing speaker recognition embeddings in TTS and provides a framework for additional future comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。