改进语音合成中口音相似度评估方法,更准更快地判断口音差异。
Pairwise Evaluation of Accent Similarity in Speech Synthesis
- 优化XAB听感测试,结合转录与差异标注,减少听众数量并提升统计效力。
- 提出基于元音共振峰和发音后验图的距离度量,有效评估口音相似性。
- 揭示常用词错误率在少数口音评估中的局限性,适合语音合成研究者参考。
尽管语音合成中高保真口音生成日益受到关注,但口音相似度的评估仍缺乏深入研究。本文旨在改进主观与客观评价方法。主观方面,通过引入转录文本、让听者标记感知到的口音差异,并实施严格可靠性筛选,优化XAB听感测试,实现更低成本与更少听众下的更高统计显著性。客观方面,采用基于元音共振峰距离与发音后验图(Phonetic Posteriorgrams)的语音相关度量,用于评估口音生成效果。对比实验表明,这些度量可与口音相似度、说话人相似度及梅尔倒谱失真(Mel Cepstral Distortion)共同使用。此外,研究强调了常见指标如词错误率(Word Error Rate)在评估非主流口音时存在显著局限性。
原文摘要 · Abstract (English)
Despite growing interest in generating high-fidelity accents, evaluating accent similarity in speech synthesis has been underexplored. We aim to enhance both subjective and objective evaluation methods for accent similarity. Subjectively, we refine the XAB listening test by adding components that achieve higher statistical significance with fewer listeners and lower costs. Our method involves providing listeners with transcriptions, having them highlight perceived accent differences, and implementing meticulous screening for reliability. Objectively, we utilise pronunciation-related metrics, based on distances between vowel formants and phonetic posteriorgrams, to evaluate accent generation. Comparative experiments reveal that these metrics, alongside accent similarity, speaker similarity, and Mel Cepstral Distortion, can be used. Moreover, our findings underscore significant limitations of common metrics like Word Error Rate in assessing underrepresented accents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。