同时预测语音质量的多个维度及其不确定性,提升故障诊断能力。
Multivariate Probabilistic Assessment of Speech Quality
- 用多变量高斯模型联合建模语音质量四大维度
- 在点估计上达顶尖水平,且给出各维度相关性与置信区间
- 适合需要精准定位语音失真问题的开发者与研究者
平均意见分(MOS)是语音质量评估的标准指标,但其单一评分无法在得分偏低时识别具体失真类型。NISQA数据集通过提供噪声、音色失真、中断感和响度四个附加维度的评分,弥补了这一不足。本文将传统的单变量MOS估计扩展为多变量框架,采用多变量高斯分布联合建模这些维度。通过使用Cholesky分解预测协方差,避免了强假设限制,并将概率仿射变换推广至多变量场景。实验表明,该模型在点估计性能上达到当前最优水平,同时首次提供各质量维度间的不确定性与相关性估计,有助于更精准诊断语音质量问题,指导针对性优化。
原文摘要 · Abstract (English)
The mean opinion score (MOS) is a standard metric for assessing speech quality, but its singular focus fails to identify specific distortions when low scores are observed. The NISQA dataset addresses this limitation by providing ratings across four additional dimensions: noisiness, coloration, discontinuity, and loudness, alongside MOS. In this paper, we extend the explored univariate MOS estimation to a multivariate framework by modeling these dimensions jointly using a multivariate Gaussian distribution. Our approach utilizes Cholesky decomposition to predict covariances without imposing restrictive assumptions and extends probabilistic affine transformations to a multivariate context. Experimental results show that our model performs on par with state-of-the-art methods in point estimation, while uniquely providing uncertainty and correlation estimates across speech quality dimensions. This enables better diagnosis of poor speech quality and informs targeted improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。