arXiv:2509.19928eess.AS2025-09中稿 · ICASSP 2026被引 10

提出新指标DS-WED,更准确衡量语音韵律多样性。

Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration

  • 用语义标记的加权编辑距离量化语音韵律变化
  • 在1000段语音上与人工评分相关性显著提升
  • 适合评估零样本语音合成系统的表达能力

韵律多样性对实现零样本文本转语音(TTS)的自然与表现力至关重要。然而,现有声学指标仅反映韵律变化的部分特征,且与人类感知相关性低,导致该问题长期缺乏可靠度量。为此,我们构建了ProsodyEval数据集,包含1000段来自7个主流TTS系统的语音样本及2000次人工评分。基于此,提出离散化语音加权编辑距离(DS-WED),通过语义标记间的加权编辑距离量化韵律差异。实验表明,DS-WED在与人类判断的相关性上显著优于现有指标,且对HuBERT和WavLM的语音分词结果保持高度鲁棒性。利用DS-WED,我们在LibriSpeech test-clean和Seed-TTS test-en上基准测试了开源SOTA TTS系统,并发现生成范式、时长控制与强化学习等因素影响韵律多样性。此外,当前大音频语言模型在捕捉韵律变化方面仍受限。音频样本可访问:https://prosodyeval.github.io。

原文摘要 · Abstract (English)

Prosody diversity is essential for achieving naturalness and expressiveness in zero-shot text-to-speech (TTS). However, frequently used acoustic metrics capture only partial views of prosodic variation and correlate poorly with human perception, leaving the problem of reliably quantifying prosody diversity underexplored. To bridge this gap, we introduce ProsodyEval, a prosody diversity assessment dataset that provides Prosody Mean Opinion Score (PMOS) alongside conventional acoustic metrics. ProsodyEval comprises 1000 speech samples derived from 7 mainstream TTS systems, with 2000 human ratings. Building on this, we propose the Discretized Speech Weighted Edit Distance (DS-WED), a new objective diversity metric that quantifies prosodic variation via weighted edit distance over semantic tokens. Experiments on ProsodyEval show that DS-WED achieves substantially higher correlation with human judgments than existing acoustic metrics, while remaining highly robust in speech tokenization from HuBERT and WavLM. Leveraging DS-WED, we benchmark state-of-the-art open-source TTS systems on LibriSpeech test-clean and Seed-TTS test-en, and further explorations uncover several factors that influence prosody diversity, including generative modeling paradigms, duration control, and reinforcement learning. Moreover, we find that current large audio language models (LALMs) remain limited in capturing prosodic variations. Audio samples are available at https://prosodyeval.github.io.

语音合成韵律评估指标设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。