系统梳理语音克隆技术,涵盖方法、评估与防滥用研究
Voice Cloning: Comprehensive Survey
- 梳理语音克隆的术语体系与核心分支
- 覆盖少样本、零样本及多语言语音合成技术
- 适合关注语音伪造防御的研究者阅读
语音克隆在数字时代迅速发展,众多研究人员与企业致力于优化相关算法以支持多样化应用。本文旨在建立语音克隆的标准化术语体系,探讨其不同变体。首先介绍声纹适配这一基础概念,随后深入分析少样本、零样本及多语言文本转语音(TTS)等方向。最后,梳理语音克隆研究中常用的评估指标与相关数据集。本综述整合现有语音克隆算法,旨在推动生成与检测技术的发展,以应对潜在滥用风险。
原文摘要 · Abstract (English)
Voice Cloning has rapidly advanced in today's digital world, with many researchers and corporations working to improve these algorithms for various applications. This article aims to establish a standardized terminology for voice cloning and explore its different variations. It will cover speaker adaptation as the fundamental concept and then delve deeper into topics such as few-shot, zero-shot, and multilingual TTS within that context. Finally, we will explore the evaluation metrics commonly used in voice cloning research and related datasets. This survey compiles the available voice cloning algorithms to encourage research toward its generation and detection to limit its misuse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。