拆解语音自然度,发现现有评估工具漏掉多种语言学错误
Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

- 将自然度分解为10个语言学维度,构建首个分维评估基准
- 4个MOS预测器只关注音质,4个音频大模型仅部分有效且依赖提示
- 研究公开数据集与代码,助力更精准的语音合成评估
自动语音合成(TTS)评估方法(如均值意见分数预测器和音频大语言模型判官)应反映人类感知,但其对听者实际感知的语音特征捕捉程度尚不明确。本文将‘自然度’解构为涵盖10个不同感知维度的语言学基础标注体系,并据此构建首个分维度元评估基准,包含860句由受训语言学家标注的语音样本。对四个MOS预测器和四个音频大语言模型判官的基准测试显示,MOS预测器仅聚焦于声学信号质量,而音频大语言模型判官表现出选择性、提示依赖性的检测能力,且在各维度间缺乏泛化性。两类方法均未能可靠捕捉语言结构化的语音错误多样性。本研究发布数据集、标注方案及评估代码,以支持更精准、可解释的TTS评估。
原文摘要 · Abstract (English)
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。