用自然语言指令实现零样本语音风格灵活控制的语音合成系统
FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions
- 基于大语言模型,通过自然语言指令和语音参考实现风格与音色分离控制
- 在多个基准上超越现有方法,能精准遵循指令并保持语音自然度
- 适合需要个性化语音生成的场景,如虚拟助手、有声书创作
本文提出FlexiVoice,一种具备零样本语音克隆能力的文本到语音合成系统,支持灵活风格控制。说话风格由自然语言指令控制,音色特征通过语音参考实现零样本输入。系统基于大语言模型构建,接收文本输入,可选地结合自然语言指令(控制风格)和语音参考(控制音色)。采用新型渐进式后训练(PPT)策略,分三阶段逐步解锁精确可控性:首先使用直接偏好优化(DPO)使模型同时准确响应指令与语音参考;接着采用多目标组相对策略优化(GRPO)实现风格指令、参考音色与文本内容的解耦;最后对指令部分进行增强优化。实验表明,FlexiVoice在多个基准上优于基线方法,展现出强解耦能力。人工评估确认其语音自然度高、控制性强且鲁棒性好。音频样例可访问 https://flexi-voice.github.io。
原文摘要 · Abstract (English)
This study proposes FlexiVoice, a text-to-speech (TTS) synthesis system capable of flexible style control with zero-shot voice cloning. The speaking style is controlled by a natural-language instruction and the voice timbre is provided by a speech reference in zero-shot manner. FlexiVoice is built with an LLM core, which takes text as input, and also takes an optional natural language instruction and an optional speech reference to control style and timbre, respectively. FlexiVoice is equipped with a novel Progressive Post-Training (PPT) scheme that progressively unlocks accurate and flexible controllability. In particular, it first employs Direct Preference Optimization (DPO) to enable FlexiVoice to accurately follow both natural language instruction and speech reference simultaneously. It then uses a multi-objective Group Relative Policy Optimization (GRPO) to disentangle style instruction, reference timbre, and textual content. Finally, it adapts instruction GRPO for more advanced instruction following. Experimental results show that FlexiVoice surpasses competing baselines and demonstrates strong capability in decoupling control factors. Human evaluations further confirm its naturalness, controllability, and robustness. Audio samples are available at https://flexi-voice.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。