用现成语音克隆模型实现高效语音匿名,保护隐私同时保持语义清晰。
Your Voice Cloning System is Secretly a Voice Anonymizer

- 利用伪说话人条件控制,复用预训练语音克隆模型实现匿名转换。
- 在7种欧洲语言上达到近最优隐私(等错误率≈0.49),语音质量更优。
- 无需重新训练,适用于多语言场景,适合需隐私保护的语音应用。
说话人匿名化旨在去除语音中的身份特征,同时保留语言内容与语音质量。我们提出将经过27,000小时语音训练的多语言语音克隆模型XTTSv2,直接用于说话人匿名化而无需再训练。核心洞察是:XTTSv2的语音克隆能力能独立于说话人身份保留韵律结构,从而通过伪说话人条件实现语音转换。我们引入一种迭代优化策略,以最大化说话人差异性与可理解性之间的调和均值,平衡隐私与实用性。在CommonVoice与Multilingual LibriSpeech数据集上,针对七种欧洲语言的评估显示,该系统达到近最优隐私水平(等错误率 ≈ 0.49),具备竞争力的可理解性,并显著优于专用匿名化基线的语音质量,且无需语言特定训练。代码已开源:https://github.com/rm00cr/coqui-tts。
原文摘要 · Abstract (English)
Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speech, for speaker anonymization without retraining. Our key insight is that XTTSv2's voice cloning capabilities preserve prosodic structure independently of speaker identity, enabling voice conversion by conditioning on a pseudo-speaker. We introduce an iterative refinement strategy that balances privacy and utility by maximizing a harmonic mean of speaker dissimilarity and intelligibility. Evaluated on seven European languages across CommonVoice and Multilingual LibriSpeech, our system achieves near-optimal privacy (EER $\approx$ 0.49), competitive intelligibility, and substantially better speech quality than dedicated anonymization baselines, while requiring no language-specific training. We release the code here: https://github.com/rm00cr/coqui-tts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。