arXiv:2508.16332cs.SDcs.AI2025-08中稿 · the IEEE Transacti…被引 15

统一可控语音生成框架,支持语音与歌唱的灵活转换与编辑。

Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation

  • 用两个统一分词器分离内容、韵律与音色,实现跨模态控制。
  • 联合训练中引入显式与隐式韵律学习,提升语音与歌唱的衔接能力。
  • 适用于语音合成、转换与编辑,适合音乐生成与语音应用开发者。

可控人声生成,尤其是歌唱等表现性领域,仍是重大挑战。本文提出 Vevo2,一个统一的可控语音与歌唱语音生成框架。为解决标注歌唱数据稀缺问题并实现灵活控制,Vevo2 引入两个音频分词器:(1) 无需乐谱的统一韵律分词器,可从语音、歌唱甚至乐器声中捕捉韵律与旋律;(2) 统一的内容-风格分词器,编码语言内容、韵律与风格,并实现音色解耦。Vevo2 包含自回归(AR)内容-风格建模阶段,实现对文本、韵律和风格的可控生成,以及基于流匹配的声学建模阶段,支持音色控制。在语音-歌唱联合训练中,提出显式与隐式韵律学习策略以弥合两者差异。此外,设计多目标后训练任务,同时优化语音可懂度与韵律相似性。实验表明,统一建模为语音与歌唱生成带来相互增益。Vevo2 在多种合成、转换与编辑任务中均表现优异,展现强大泛化能力与实用性。音频样例见 https://versasinger.github.io/。

原文摘要 · Abstract (English)

Controllable human voice generation, particularly for expressive domains like singing, remains a significant challenge. This paper introduces Vevo2, a unified framework for controllable speech and singing voice generation. To tackle issues like the scarcity of annotated singing data and to enable flexible controllability, Vevo2 introduces two audio tokenizers: (1) a unified music-notation-free prosody tokenizer that captures prosody and melody from speech, singing, and even instrumental sounds, and (2) a unified content-style tokenizer that encodes linguistic content, prosody, and style for both speech and singing, while enabling timbre disentanglement. Vevo2 consists of an auto-regressive (AR) content-style modeling stage, which aims to enable controllability over text, prosody, and style, as well as a flow-matching acoustic modeling stage that allows for timbre control. Particularly, during the speech-singing joint training of the AR model, we propose both explicit and implicit prosody learning strategies to bridge speech and singing voice. Moreover, to further enhance the Vevo2's ability to follow text and prosody, we design a multi-objective post-training task that integrates both intelligibility and prosody similarity alignment. Experimental results show that the unified modeling in Vevo2 brings mutual benefits to both speech and singing voice generation. Additionally, Vevo2's effectiveness across a wide range of synthesis, conversion, and editing tasks for both speech and singing further demonstrates its strong generalization ability and versatility. Audio samples are are available at https://versasinger.github.io/.

语音生成歌唱合成可控生成音色解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。