统一框架实现语音生成与风格控制,高效且灵活。
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation

- 将语音分解为内容、语调、风格三要素,统一建模
- 仅4351万参数,4步ODE推理,生成质量媲美大模型
- 支持语音合成、变声、编辑等多任务,风格控制更精准
人类语音生成在语音合成、歌唱语音生成、语音克隆和语音编辑方面取得了快速进展。然而,现有系统大多针对特定任务设计,依赖任务相关的架构、控制信号或自回归解码,限制了细粒度控制能力和推理效率。本文提出CookVoice,一个统一的多模态、多风格、多任务人类语音生成框架。CookVoice将人声分解为内容、语调和风格三个关键因素,实现在统一模型中同时支持语音和歌唱语音生成。为实现精确灵活的控制,设计了一种灵活对齐策略,将文本、风格和语调控制信号映射到频谱图的帧级。该设计使CookVoice可支持多种任务,包括文本转语音、文本转歌唱语音、风格可控生成、语音模仿、语音转换和语音编辑。实验表明,CookVoice生成质量与现有文本转语音及文本转歌唱语音基线相当,同时具备更强的风格和语调控制能力。此外,仅用4351万参数和最少4步常微分方程(ODE)步骤即可达到与大规模基线相当的性能,适用于实际语音生成应用。演示页面见:https://haoweilou.github.io/CookVoice/
原文摘要 · Abstract (English)
Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. However, most existing systems are designed for specific tasks and often rely on task-dependent architectures, control signals, or autoregressive decoding, limiting fine-grained controllability and inference efficiency. In this paper, we propose CookVoice, a unified framework for multimodal, multi-style, and multi-task human voice generation. CookVoice decomposes the human voice into three key factors: content, prosody, and style, enabling both speech and singing voice generation within a unified model. To achieve precise and flexible controllability, we design a flexible alignment strategy that maps text, style, and prosody control signals onto the frame-level of spectrogram. This design allows CookVoice to support a wide range of tasks, including text-to-speech, text-to-singing voice, style-controllable generation, voice mimicry, voice conversion, and voice editing. Experimental results show that CookVoice achieves generation quality comparable to existing Text-to-Speech and text-to-singing voice baselines, while providing stronger style and prosody controllability. Moreover, CookVoice achieves comparable performance to large-scale baselines with only 43.51 million parameters and efficient inference using as few as 4 ODE steps, making it a practical solution for real-world human voice generation applications. Demo page is available at https://haoweilou.github.io/CookVoice/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。