轻量级零样本语音合成,高效分离语音内容与说话人特征
Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement
- 通过两阶段自蒸馏框架,从训练数据角度解耦语音内容与说话人信息
- 在零样本语音合成任务中表现优异,推理时延仅0.13(CPU)和0.012(GPU)RTF
- 模型轻量稳定,适合资源受限场景下的个性化语音生成应用
零样本文本到语音合成在通过语音克隆实现个性化语音定制方面展现出巨大潜力。然而,当前方法严重依赖大模型规模和大规模训练数据以保证性能与跨说话人的泛化能力,带来部署成本高和数据安全风险。本文提出一种轻量且稳定的零样本语音合成系统。设计新架构,分别从源语音和提示语音中建模语言内容与说话人属性。提出两阶段自蒸馏框架,从训练数据视角构建并行数据对,有效解耦语言内容与说话人信息。大量实验表明,该系统在零样本语音合成任务中表现出色,具备卓越稳定性。同时计算效率显著提升,CPU与GPU上的实时因子(RTF)分别为0.13和0.012。
原文摘要 · Abstract (English)
Zero-shot Text-To-Speech (TTS) synthesis shows great promise for personalized voice customization through voice cloning. However, current methods for achieving zero-shot TTS heavily rely on large model scales and extensive training datasets to ensure satisfactory performance and generalizability across various speakers. This raises concerns regarding both deployment costs and data security. In this paper, we present a lightweight and stable zero-shot TTS system. We introduce a novel TTS architecture designed to effectively model linguistic content and various speaker attributes from source speech and prompt speech, respectively. Furthermore, we present a two-stage self-distillation framework that constructs parallel data pairs for effectively disentangling linguistic content and speakers from the perspective of training data. Extensive experiments show that our system exhibits excellent performance and superior stability on the zero-shot TTS tasks. Moreover, it shows markedly superior computational efficiency, with RTFs of 0.13 and 0.012 on the CPU and GPU, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。