让语音与歌唱自动切换,仅凭文本内容就实现无缝融合。
UniVocal: Unified Speech-Singing Code-Switching Synthesis

- 通过文本上下文隐式推断演唱模式,无需额外控制标签。
- 在SCSBench上达到顶尖性能,同时保持普通语音和歌唱生成效果。
- 提出新数据合成管道和多场景评测基准,适合语音合成研究者。
我们提出UniVocal,一个统一框架,通过文本上下文隐式推断演唱模式,开创语音-歌唱混用合成(SCS)任务——即转换由文本语义自主驱动,类似人类语言自然融合。不同于单模式生成或依赖切换标签的系统,UniVocal仅依据文本上下文推断演唱模式。为此,我们采用高效两阶段课程学习策略,逐步训练出具备所需SCS能力的高质量语音合成系统。针对数据稀缺问题,我们构建可扩展的数据合成流程,生成在语义与声学上均自然的多样化混用数据,并推出新多场景评测基准SCSBench。为克服语义分词器对声学细节捕捉不足的问题,我们引入精细化音节标记和思维链(CoT)生成机制,在内容生成前规划韵律,显著提升共情式语音与歌唱旋律质量。实验表明,UniVocal在SCSBench上表现领先,同时在常规语音与歌唱任务中保持竞争力。音频样例见https://project-univocal-demo.github.io/demo/,代码与数据集已开源于https://github.com/FunAudioLLM/FunResearch/tree/main/UniVocal。
原文摘要 · Abstract (English)
We propose UniVocal, a unified framework that implicitly infers vocal modes from text context to pioneer Speech-Singing Code-Switching (SCS) Synthesis - a task where transitions are autonomously driven by textual semantics, akin to seamless human language blending. Unlike single-mode generation or systems relying on switching-control tags, our proposed UniVocal implicitly infers vocal modes solely from text context. To achieve this, we employ a data-efficient two-stage curriculum learning strategy that progressively trains a competitive TTS system to acquire the desired SCS capability. Addressing data scarcity, we introduce a scalable pipeline to synthesize diverse code-switching data that is both semantically and acoustically natural, alongside a new multi-scenario benchmark, SCSBench. To address limitations of semantic tokenizers in capturing acoustic details, we also introduce refined cent token and Chain-of-Thought (CoT) generation for planning prosody before content generation, effectively enhancing empathetic speech generation and singing melody. Experimental results demonstrate that UniVocal achieves state-of-the-art performance on SCSBench while maintaining competitive performance on regular speech and singing tasks. Audio samples are available at https://project-univocal-demo.github.io/demo/. The code and dataset are released at https://github.com/FunAudioLLM/FunResearch/tree/main/UniVocal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。