对比多种语音评估方法,发现WavLM基线特征最贴近人耳判断。
Understanding Frechet Speech Distance for Synthetic Speech Quality Evaluation
- 用不同语音嵌入和设置测试弗雷歇语音距离与SMMD
- WavLM Base+特征与人工评分相关性最强
- 适合大规模语音合成质量评估,替代部分主观测试
合成语音的客观评估仍是关键挑战。人工听辨是金标准,但成本高、难以规模化。弗雷歇距离成为有前景的替代方案,但其可靠性高度依赖嵌入选择与实验设置。本文全面评估了弗雷歇语音距离(FSD)及其变体语音最大均值差异(SMMD),在多种嵌入和条件下表现。同时结合人工听辨、语音可懂度及合成训练的ASR WER验证这些指标的感知相关性。结果表明,WavLM Base+特征与人工评分具有最稳定的对应关系。尽管FSD与SMMD无法完全替代主观评价,但可作为低成本、可复现的补充工具,尤其适用于大规模或无法直接听辨的场景。代码已开源:https://github.com/kaen2891/FrechetSpeechDistance。
原文摘要 · Abstract (English)
Objective evaluation of synthetic speech quality remains a critical challenge. Human listening tests are the gold standard, but costly and impractical at scale. Fréchet Distance has emerged as a promising alternative, yet its reliability depends heavily on the choice of embeddings and experimental settings. In this work, we comprehensively evaluate Fréchet Speech Distance (FSD) and its variant Speech Maximum Mean Discrepancy (SMMD) under varied embeddings and conditions. We further incorporate human listening evaluations alongside TTS intelligibility and synthetic-trained ASR WER to validate the perceptual relevance of these metrics. Our findings show that WavLM Base+ features yield the most stable alignment with human ratings. While FSD and SMMD cannot fully replace subjective evaluation, we show that they can serve as complementary, cost-efficient, and reproducible measures, particularly useful when large-scale or direct listening assessments are infeasible. Code is available at https://github.com/kaen2891/FrechetSpeechDistance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。