arXiv:2412.04724eess.AScs.SD2024-12中稿 · AAAI被引 26

让语音转换更灵活快速,能独立控制音色与风格。

StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

  • 用条件流匹配分解语音,分别控制音色和风格。
  • 生成速度比自回归方法快25倍,比扩散模型快1.65倍。
  • 无需训练即可适配新说话人,适合需要高效定制语音的场景。

零样本语音转换(VC)旨在将源说话人的音色迁移到任意未见目标说话人,同时保持原始语言内容。尽管基于语言模型或扩散模型的零样本VC取得进展,仍存在三大挑战:1)现有方法难以独立迁移音色与风格至不同未见说话人;2)依赖自回归建模或大量采样步骤,推理速度慢;3)转换语音的质量与相似性仍不理想。为此,本文提出名为StableVC的可控零样本语音转换方法,可独立迁移音色与风格至不同未见说话人。具体地,将语音分解为语言内容、音色与风格三部分,通过条件流匹配模块重建高质量梅尔频谱图。为有效在零样本下捕捉音色与风格,设计新颖的双注意力机制与自适应门控结构,替代传统特征拼接。非自回归设计使StableVC能高效捕获复杂音色与风格,生成速度显著超过实时。实验表明,StableVC优于现有最先进系统,在音色与风格控制上更具灵活性;且相比自回归基线提速约25倍,相比扩散基线提速1.65倍。

原文摘要 · Abstract (English)

Zero-shot voice conversion (VC) aims to transfer the timbre from the source speaker to an arbitrary unseen speaker while preserving the original linguistic content. Despite recent advancements in zero-shot VC using language model-based or diffusion-based approaches, several challenges remain: 1) current approaches primarily focus on adapting timbre from unseen speakers and are unable to transfer style and timbre to different unseen speakers independently; 2) these approaches often suffer from slower inference speeds due to the autoregressive modeling methods or the need for numerous sampling steps; 3) the quality and similarity of the converted samples are still not fully satisfactory. To address these challenges, we propose a style controllable zero-shot VC approach named StableVC, which aims to transfer timbre and style from source speech to different unseen target speakers. Specifically, we decompose speech into linguistic content, timbre, and style, and then employ a conditional flow matching module to reconstruct the high-quality mel-spectrogram based on these decomposed features. To effectively capture timbre and style in a zero-shot manner, we introduce a novel dual attention mechanism with an adaptive gate, rather than using conventional feature concatenation. With this non-autoregressive design, StableVC can efficiently capture the intricate timbre and style from different unseen speakers and generate high-quality speech significantly faster than real-time. Experiments demonstrate that our proposed StableVC outperforms state-of-the-art baseline systems in zero-shot VC and achieves flexible control over timbre and style from different unseen speakers. Moreover, StableVC offers approximately 25x and 1.65x faster sampling compared to autoregressive and diffusion-based baselines.

语音转换风格控制流匹配零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。