BoN语音合成评估受语音识别模型家族影响,跨家族评估更可靠
Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment
- 用不同语音识别器评估同一语音合成结果,排名差异显著
- 跨家族评估比同家族评估更能接近理想性能,提升12%词错误率表现
- 提出融合多个识别器排名的新方法,适合追求公平评估的研究者
Best-of-$N$(BoN)推理通过自动语音识别(ASR)验证器从多个候选输出中选择内容一致的语音,提升零样本文本到语音合成质量。我们发现评估存在混淆:验证器表现显著依赖于所用ASR家族。在LibriSpeech-PC数据集上使用F5-TTS时,Whisper、wav2vec 2.0和HuBERT等不同家族的评价器导致验证器排名差异显著;而相同家族的验证器与评价器组合比跨家族组合恢复更多理想性能头距,尽管其表示高度相似。这表明评估偏差源于身份或谱系关联,而非泛化表征相似性。为缓解此偏见,我们提出两种跨家族排名集成策略:排名平均与合取最大排名。两者均在独立评价器上降低平均词错误率(WER),且不损害自动相似度或质量指标,最优集成在N=10时相较F5-TTS实现12%相对WER降低。这些发现支持以多评价器三角验证作为报告BoN语音合成性能的更可靠默认方式。
原文摘要 · Abstract (English)
Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting among multiple candidates with an automatic speech recognition (ASR) verifier. We identify an evaluation confound: the apparent quality of a verifier depends strongly on the ASR family used for evaluation. On LibriSpeech-PC with F5-TTS, verifier rankings vary substantially across Whisper, wav2vec 2.0, and HuBERT evaluators, while same-family verifier and evaluator pairs recover considerably more oracle headroom than cross-family pairs despite highly similar representations. This pattern suggests identity- or lineage-level coupling rather than general representational similarity. To mitigate this bias, we propose two cross-family rank ensembles: rank averaging and conjunctive max-rank. Both improve mean word error rate across independent evaluators without degrading automatic similarity or quality metrics, and the best ensemble achieves a $12\%$ relative WER reduction over F5-TTS at $N=10$. These findings motivate cross-evaluator triangulation as a more reliable default for reporting BoN TTS performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。