让语音合成听懂自然语言指令,实现更灵活的风格控制。
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
- 用推理机制将自然语言指令转化为声学特征
- 在新数据集上指令遵循准确率提升显著
- 适合内容创作者实现个性化语音生成
指令式语音合成(InstructTTS)利用自然语言描述作为风格提示来引导语音生成。然而现有方法主要依赖音频相关标签或其变体表述,难以处理灵活、高层次的指令。这种僵化控制无法满足内容创作者通过描述性指令引导生成的需求。为此,我们提出OV-InstructTTS,一种全新的开放词汇指令语音合成范式。该方案包含自建数据集OV-Speech和新型推理驱动框架。OV-Speech数据集将语音与开放词汇指令配对,并为每条指令附加连接高层指令与声学特征的推理过程。推理驱动框架从开放词汇指令中推断情感、声学及副语言信息后进行语音合成。评估表明,该方法显著提升了指令遵循保真度与语音表现力。我们相信此项工作将推动更具泛化能力与实际应用价值的下一代用户友好型语音合成系统发展。数据集与演示已公开于项目主页。
原文摘要 · Abstract (English)
Instruct Text-to-Speech (InstructTTS) leverages natural language descriptions as style prompts to guide speech synthesis. However, existing InstructTTS methods mainly rely on a direct combination of audio-related labels or their diverse rephrasings, making it difficult to handle flexible, high-level instructions. Such rigid control is insufficient for users such as content creators who wish to steer generation with descriptive instructions. To address these constraints, we introduce OV-InstructTTS, a new paradigm for open-vocabulary InstructTTS. We propose a comprehensive solution comprising a newly curated dataset, OV-Speech, and a novel reasoning-driven framework. The OV-Speech dataset pairs speech with open-vocabulary instructions, each augmented with a reasoning process that connects high-level instructions to acoustic features. The reasoning-driven framework infers emotional, acoustic, and paralinguistic information from open-vocabulary instructions before synthesizing speech. Evaluations show that this reasoning-driven approach significantly improves instruction-following fidelity and speech expressiveness. We believe this work can inspire the next user-friendly InstructTTS systems with stronger generalization and real-world applicability. The dataset and demos are publicly available on our project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。