让语音转换更自然,还能自由控制说话节奏。
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
- 用内容令牌和数据扰动减少无关信息干扰
- 仅需两步采样即可生成高保真语音,延迟极低
- 可精准迁移目标说话者的语速节奏,适合音色克隆
零样本语音转换旨在将源说话人的音色转换为任意未见的目标音色,同时保留语音内容。现有方法多关注保持源语音的韵律,但细粒度音色信息可能通过韵律泄露,且目标韵律的迁移很少被研究。为此,我们提出R-VC,一种节奏可控且高效的零样本语音转换模型。R-VC采用数据扰动技术,并将源语音离散化为Hubert内容令牌,消除大量与内容无关的信息。通过引入掩码生成式变压器进行上下文持续时间建模,模型能将语言内容的时长适配至期望的目标语速风格,实现目标说话人节奏的迁移。此外,R-VC在训练中引入具有捷径流匹配的扩散变换器(DiT),使网络不仅依赖当前噪声水平,还基于期望的采样步长进行条件控制,从而在极少采样步数(甚至仅两步)下生成高音色相似性和高质量语音,显著降低延迟。实验表明,R-VC在更小数据集上达到与顶尖方法相当的说话人相似度,且在语音自然度、可懂度和风格迁移性能上均超越现有方法。
原文摘要 · Abstract (English)
Zero-Shot Voice Conversion (VC) aims to transform the source speaker's timbre into an arbitrary unseen one while retaining speech content. Most prior work focuses on preserving the source's prosody, while fine-grained timbre information may leak through prosody, and transferring target prosody to synthesized speech is rarely studied. In light of this, we propose R-VC, a rhythm-controllable and efficient zero-shot voice conversion model. R-VC employs data perturbation techniques and discretize source speech into Hubert content tokens, eliminating much content-irrelevant information. By leveraging a Mask Generative Transformer for in-context duration modeling, our model adapts the linguistic content duration to the desired target speaking style, facilitating the transfer of the target speaker's rhythm. Furthermore, R-VC introduces a powerful Diffusion Transformer (DiT) with shortcut flow matching during training, conditioning the network not only on the current noise level but also on the desired step size, enabling high timbre similarity and quality speech generation in fewer sampling steps, even in just two, thus minimizing latency. Experimental results show that R-VC achieves comparable speaker similarity to state-of-the-art VC methods with a smaller dataset, and surpasses them in terms of speech naturalness, intelligibility and style transfer performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。