arXiv:2505.23009cs.LGcs.SD2025-05NeurIPS被引 34

用AI自动生成复杂语音测试用例,精准评估语音合成模型表现

EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge

  • 用大模型迭代生成包含情感、语法等6类挑战的1645个测试用例
  • 通过音频大模型评分,与人工评价高度一致,能发现细微性能差异
  • 适合研究者和工程师测试真实场景下的语音合成能力

文本到语音(TTS)基准测试常难以捕捉模型处理语义复杂文本的能力。基于EmergentTTS,我们提出EmergentTTS-Eval,一个涵盖六类挑战性场景的综合性基准:情感表达、副语言特征、外语词、句法复杂度、复杂发音(如网址、公式)及疑问句。关键在于,该框架自动化生成测试用例并进行评估,可轻松扩展。从少量人工编写的种子提示出发,利用大模型迭代生成针对特定结构、语音和语调挑战的样本,共生成1,645个多样化测试用例。此外,采用模型作为裁判的方法,使用大型音频语言模型(LALM)从情感表达、语调、语调和发音准确性等多个维度评估语音质量。我们在11Labs、Deepgram、OpenAI 4o-mini-TTS等先进开源与专有TTS系统上进行了评估,验证了其揭示细粒度性能差异的能力。结果表明,模型作为裁判的方法具有强鲁棒性,且与人类偏好高度相关。代码与数据集已开源。

原文摘要 · Abstract (English)

Text-to-Speech (TTS) benchmarks often fail to capture how well models handle nuanced and semantically complex text. Building on $\textit{EmergentTTS}$, we introduce $\textit{EmergentTTS-Eval}$, a comprehensive benchmark covering six challenging TTS scenarios: emotions, paralinguistics, foreign words, syntactic complexity, complex pronunciation (e.g. URLs, formulas), and questions. Crucially, our framework automates both test-case generation and evaluation, making the benchmark easily extensible. Starting from a small set of human-written seed prompts, we iteratively extend them using LLMs to target specific structural, phonetic and prosodic challenges, resulting in 1,645 diverse test cases. Moreover, we employ a model-as-a-judge approach, using a Large Audio Language Model (LALM) to assess the speech across multiple dimensions such as expressed emotion, prosodic, intonational, and pronunciation accuracy. We evaluate state-of-the-art open-source and proprietary TTS systems, such as 11Labs, Deepgram, and OpenAI's 4o-mini-TTS, on EmergentTTS-Eval, demonstrating its ability to reveal fine-grained performance differences. Results show that the model-as-a-judge approach offers robust TTS assessment and a high correlation with human preferences. We open source the evaluation $\href{https://github.com/boson-ai/EmergentTTS-Eval-public}{code}$ and the $\href{https://huggingface.co/datasets/bosonai/EmergentTTS-Eval}{dataset}$.

语音合成自动评测大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。