arXiv:2603.24430cs.SD2026-03

通过迭代合成提升零样本语音合成评估的区分度与可靠性

Iterate to Differentiate: Enhancing Discriminability and Reliability in Zero-Shot TTS Evaluation

  • 用模型自产语音作为参考,递归合成以制造分布偏移
  • 表现更好的模型在迭代中衰减更慢,差距被放大
  • 使客观指标更贴近人评,适合评测先进零样本语音模型

现代零样本文本到语音(TTS)模型的可靠评估仍具挑战性。主观评测成本高且难以复现,而客观指标常出现饱和,无法区分最先进系统。为此,我们提出迭代区分(I2D)评估框架,通过递归使用模型自身输出作为参考来合成语音。高质量模型对迭代合成带来的分布偏移更具鲁棒性,性能下降更缓慢。I2D利用这种差异性退化来放大性能差距并揭示模型稳健性。通过聚合多轮迭代的客观指标,I2D提升了区分能力,使系统级UTMOSv2的SRCC从0.118提升至0.464。在中文、英文及情感数据集上对11个模型的实验表明,I2D能实现更可靠的自动化评估。

原文摘要 · Abstract (English)

Reliable evaluation of modern zero-shot text-to-speech (TTS) models remains challenging. Subjective tests are costly and hard to reproduce, while objective metrics often saturate, failing to distinguish SOTA systems. To address this, we propose Iterate to Differentiate (I2D), an evaluation framework that recursively synthesizes speech using the model's own outputs as references. Higher-quality models exhibit greater resilience to the distributional shift induced by iterative synthesis, resulting in slower performance degradation. I2D exploits this differential degradation to amplify performance gaps and reveal robustness. By aggregating objective metrics across iterations, I2D improves discriminability and alignment with human judgments, increasing system-level SRCC from 0.118 to 0.464 for UTMOSv2. Experiments on 11 models across Chinese, English, and emotion datasets demonstrate that I2D enables more reliable automated evaluation for zero-shot TTS.

语音合成评估方法零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。