对10种印度语言的TTS系统进行大规模听觉偏好评估,揭示语音质量与表达力的权衡。
Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

- 设计多维度对比评估框架,控制语言多样性并捕捉真实听感
- 收集超12万次对比判断,覆盖5000+句子和7个先进TTS模型
- 发现语音清晰度与自然度存在权衡,适合多语言语音系统研究者
众包式成对评估已成为评估基础模型的可扩展方法。然而,将其应用于文本转语音(TTS)时,由于语言多样性及语音感知的多维性,导致评估结果方差较大。本文提出一种受控的多维成对评估框架,结合语言控制与感知基准标注。基于10种印地语系语言中超过5000个原生及代码混杂句子,评估7个领先TTS系统,并从1900多名母语评价者处收集了超过12万次成对比较。除总体偏好外,评价者还针对6个感知维度打分:可懂性、表现力、音质、生动性、噪声和幻觉。通过布拉德利-特里模型构建多语言排行榜,利用SHAP分析解释人类偏好,并分析排行榜可靠性以及模型在各感知维度上的优劣与权衡。
原文摘要 · Abstract (English)
Crowdsourced pairwise evaluation has emerged as a scalable approach for assessing foundation models. However, applying it to Text to Speech(TTS) introduces high variance due to linguistic diversity and multidimensional nature of speech perception. We present a controlled multidimensional pairwise evaluation framework for multilingual TTS that combines linguistic control with perceptually grounded annotation. Using 5K+ native and code-mixed sentences across 10 Indic languages, we evaluate 7 state-of-the-art TTS systems and collect over 120K pairwise comparisons from over 1900 native raters. In addition to overall preference, raters provide judgments across 6 perceptual dimensions: intelligibility, expressiveness, voice quality, liveliness, noise, and hallucinations. Using Bradley-Terry modeling, we construct a multilingual leaderboard, interpret human preference using SHAP analysis and analyze leaderboard reliability alongside model strengths and trade-offs across perceptual dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。