arXiv:2502.07243cs.SDcs.AI2025-02ICLR被引 78

无需标注数据,可零样本控制音色与语调的语音模仿新框架。

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

  • 用自监督方法解耦语音内容、音色与风格,实现可控生成。
  • 仅用60K小时音频训练,零样本下仍媲美或超越现有方法。
  • 适合语音合成、语音转换等需要高可控性的场景。

语音模仿对音色和语调等特定属性至关重要,但现有方法依赖标注数据,难以有效解耦音色与风格,尤其在零样本场景下控制能力差。为此,我们提出Vevo,一种无需标注数据的零样本语音模仿框架。该框架分两阶段:(1) 内容-风格建模:输入文本或语音内容标记,通过自回归Transformer生成内容-风格标记,由风格参考引导;(2) 声学建模:输入内容-风格标记,使用流匹配Transformer生成声学表示,由音色参考引导。为获取语音的内容与内容-风格标记,设计全自监督方法,逐步解耦音色、风格与语言内容。采用HuBERT的连续隐藏特征作为输入,以VQ-VAE为编码器,通过调节码本词汇量作为信息瓶颈,获得解耦表征。模型仅在60K小时有声书数据上自监督训练,未在特定风格语料上微调,却在口音和情感转换任务中表现媲美或优于现有方法。此外,其在零样本语音转换与文本到语音任务中的优异表现,验证了强泛化性与通用性。音频样例见https://versavoice.github.io。

原文摘要 · Abstract (English)

The imitation of voice, targeted on specific speech attributes such as timbre and speaking style, is crucial in speech generation. However, existing methods rely heavily on annotated data, and struggle with effectively disentangling timbre and style, leading to challenges in achieving controllable generation, especially in zero-shot scenarios. To address these issues, we propose Vevo, a versatile zero-shot voice imitation framework with controllable timbre and style. Vevo operates in two core stages: (1) Content-Style Modeling: Given either text or speech's content tokens as input, we utilize an autoregressive transformer to generate the content-style tokens, which is prompted by a style reference; (2) Acoustic Modeling: Given the content-style tokens as input, we employ a flow-matching transformer to produce acoustic representations, which is prompted by a timbre reference. To obtain the content and content-style tokens of speech, we design a fully self-supervised approach that progressively decouples the timbre, style, and linguistic content of speech. Specifically, we adopt VQ-VAE as the tokenizer for the continuous hidden features of HuBERT. We treat the vocabulary size of the VQ-VAE codebook as the information bottleneck, and adjust it carefully to obtain the disentangled speech representations. Solely self-supervised trained on 60K hours of audiobook speech data, without any fine-tuning on style-specific corpora, Vevo matches or surpasses existing methods in accent and emotion conversion tasks. Additionally, Vevo's effectiveness in zero-shot voice conversion and text-to-speech tasks further demonstrates its strong generalization and versatility. Audio samples are available at https://versavoice.github.io.

语音模仿自监督零样本音色控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。