arXiv:2509.20485eess.AScs.LG2025-09中稿 · IEEE Open Journal …被引 1

提出新评估框架TTScore,精准衡量语音合成的可懂度与语调质量。

Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens

  • 基于离散语音标记的条件预测,实现无需参考音频的评估。
  • 在SOMOS、VoiceMOS等数据集上与人工评价相关性更高。
  • 适合语音合成研究者用于细化评估模型性能。

客观评估合成语音对推进语音生成系统至关重要,但现有可懂度和语调评估指标范围有限且与人类感知相关性弱。词错误率(WER)仅提供粗略的文本级可懂度度量,而F0-RMSE等基音指标则局限于参考依赖的狭窄视角。为此,我们提出TTScore,一种基于离散语音标记条件预测的目标导向、无参考评估框架。TTScore包含两个序列到序列预测器:TTScore-int通过内容标记衡量可懂度,TTScore-pro通过语调标记评估语调。对每段合成语音,预测器计算对应标记序列的似然值,生成可解释的得分,反映与预期语言内容和语调结构的对齐程度。在SOMOS、VoiceMOS和TTSArena基准上的实验表明,TTScore-int与TTScore-pro提供可靠、针对性的评估,并在整体质量的人工评价相关性上优于现有指标。

原文摘要 · Abstract (English)

Objective evaluation of synthesized speech is critical for advancing speech generation systems, yet existing metrics for intelligibility and prosody remain limited in scope and weakly correlated with human perception. Word Error Rate (WER) provides only a coarse text-based measure of intelligibility, while F0-RMSE and related pitch-based metrics offer a narrow, reference-dependent view of prosody. To address these limitations, we propose TTScore, a targeted and reference-free evaluation framework based on conditional prediction of discrete speech tokens. TTScore employs two sequence-to-sequence predictors conditioned on input text: TTScore-int, which measures intelligibility through content tokens, and TTScore-pro, which evaluates prosody through prosody tokens. For each synthesized utterance, the predictors compute the likelihood of the corresponding token sequences, yielding interpretable scores that capture alignment with intended linguistic content and prosodic structure. Experiments on the SOMOS, VoiceMOS, and TTSArena benchmarks demonstrate that TTScore-int and TTScore-pro provide reliable, aspect-specific evaluation and achieve stronger correlations with human judgments of overall quality than existing intelligibility and prosody-focused metrics.

语音合成评估方法可懂度语调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。