arXiv:2606.19951eess.AScs.CL2026-06中稿 · INTERSPEECH 2026

模型对语音质量的判断,远不如人耳敏感。

Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

  • 通过声学与语调扰动测试模型与人类感知差异
  • 模型能识别声学退化但忽略语调错误,人类评分下降显著
  • 模型对音高均值敏感却无视语速和音高变化,人耳则能察觉

均值意见分(MOS)预测模型广泛用作文本到语音(TTS)研究中的代理指标,但其是否能捕捉超越声学保真度的质量差异尚不明确。本文通过控制语音的声学退化、语调错误以及说话人特有特征(如基频和语速)的扰动,获取了人类听者与模型对这些语音样本的MOS预测,并分析了两者感知特性的差异。结果显示,大多数模型能较好追踪声学退化,但对语调错误均不敏感,尽管人类评分出现大幅下降。对于说话人特征,模型表现出双重分离:对平均基频(F0)存在明显偏差,而人类评分中并无此现象;但对语速和F0变异性则不敏感,人类却能察觉。这些发现揭示了标量MOS预测在声学保真度之外的局限性。

原文摘要 · Abstract (English)

Mean opinion score (MOS) prediction models are widely used as proxy metrics in text-to-speech (TTS) research, yet their ability to capture quality differences beyond acoustic fidelity remains unclear. We investigate this via controlled perturbations on speech: acoustic degradation, prosodic errors, and manipulation of speaker-specific characteristics such as pitch and speaking rate. We obtained MOS predictions for these speech samples from both human listeners and the model, and analyzed the differences in their perceptual characteristics. Results show that most models track acoustic degradation well, while all are insensitive to prosodic errors despite large subjective score drops. For speaker characteristics, models exhibit a double dissociation: strong mean fundamental frequency (F0) biases absent in human ratings, yet insensitivity to speaking rate and F0 variability that humans notice. These findings highlight limitations of scalar MOS prediction beyond acoustic fidelity.

语音评估声学模型人类感知语调分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。