构建首个复杂自然语言指令跟随的语音合成评测基准,提升语音合成灵活性。
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
- 设计三类指令任务,覆盖参数指定、风格描述与角色扮演。
- 含6000个中英文测试用例,用Gemini自动评估指令遵循能力。
- 发现现有系统仍有巨大改进空间,适合语音合成研究者使用。
在现代语音合成中,韵律信息如音色、情感状态和动态语调对传达语义之外的细微差别至关重要。传统文本到语音(TTS)系统依赖固定风格标签或插入语音提示来控制这些特征,严重限制了灵活性。近期尝试采用自然语言指令调节韵律特征,显著提升了指令驱动型TTS模型的泛化能力。尽管许多TTS系统现已支持通过文字描述进行定制化合成,但其实际解析与执行复杂指令的能力仍缺乏系统评估。此外,针对指令式TTS的高质量基准和自动化评估指标仍严重不足,制约了模型的准确评估与迭代优化。为此,我们提出InstructTTSEval,一个用于衡量复杂自然语言风格控制能力的基准。包含三项任务:声学参数指定、描述性风格指令和角色扮演,每项均含中英文子集,共6000个测试案例,配以参考音频。我们采用Gemini作为自动评判器,评估模型的指令遵循能力。对现有可访问指令跟随TTS系统的评估显示,仍有巨大改进空间。我们预期InstructTTSEval将推动更强大、灵活、精准的指令跟随语音合成发展。
原文摘要 · Abstract (English)
In modern speech synthesis, paralinguistic information--such as a speaker's vocal timbre, emotional state, and dynamic prosody--plays a critical role in conveying nuance beyond mere semantics. Traditional Text-to-Speech (TTS) systems rely on fixed style labels or inserting a speech prompt to control these cues, which severely limits flexibility. Recent attempts seek to employ natural-language instructions to modulate paralinguistic features, substantially improving the generalization of instruction-driven TTS models. Although many TTS systems now support customized synthesis via textual description, their actual ability to interpret and execute complex instructions remains largely unexplored. In addition, there is still a shortage of high-quality benchmarks and automated evaluation metrics specifically designed for instruction-based TTS, which hinders accurate assessment and iterative optimization of these models. To address these limitations, we introduce InstructTTSEval, a benchmark for measuring the capability of complex natural-language style control. We introduce three tasks, namely Acoustic-Parameter Specification, Descriptive-Style Directive, and Role-Play, including English and Chinese subsets, each with 1k test cases (6k in total) paired with reference audio. We leverage Gemini as an automatic judge to assess their instruction-following abilities. Our evaluation of accessible instruction-following TTS systems highlights substantial room for further improvement. We anticipate that InstructTTSEval will drive progress toward more powerful, flexible, and accurate instruction-following TTS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。