arXiv:2609.03283cs.SD2026-09

语音质量评估需兼顾语义与声学细节,仅靠语义不够

Is Semantics Enough for Speech Mean Opinion Score Prediction?

论文配图:Is Semantics Enough for Speech Mean Opinion Score Prediction?
图 1 · 摘自论文原文
  • 融合语义与声学建模的统一编码器表现更优
  • 在标准和域外数据集上均达到更高预测上限
  • 适合语音合成质量评估与模型优化研究者

平均意见分(MOS)是评估合成语音自然度的黄金标准。然而,当前自动MOS预测模型主要依赖自监督学习(SSL)模型,侧重高层语义,可能削弱对关键声学细节的捕捉能力。本文系统比较了三种表征范式:SSL、纯声学神经音频编解码器(NAC)以及将语义融入重建架构的统一NAC。在标准BVCC及多个域外(OOD)数据集上的广泛评测表明,结合语义理解与细粒度声学建模的特征能实现语音质量评估的更高性能上限。最终结果表明,仅依赖语义不足以实现鲁棒的MOS预测,必须同时关注语义内容与声学保真度。

原文摘要 · Abstract (English)

Mean Opinion Score (MOS) is the gold standard for evaluating synthesized speech naturalness. However, current automatic MOS predictors are dominated by self-supervised learning (SSL) models that prioritize high-level semantics, potentially compromising their ability to capture critical acoustic details. In this paper, we systematically investigate representations from three paradigms: SSLs, acoustic-only neural audio codecs (NACs), and unified NACs that integrate semantics into reconstruction-based architectures. Extensive benchmarking on the standard BVCC and multiple out-of-domain (OOD) datasets demonstrates that features synergizing semantic understanding with fine-grained acoustic modeling achieve a higher performance upper bound in speech quality assessment. Ultimately, our findings highlight that semantics alone are not enough; a dual focus on semantic content and acoustic fidelity is essential for robust MOS prediction.

语音评估声学建模语义融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。