arXiv:2410.17834eess.AScs.LG2024-10中稿 · Interspeech 2025被引 5

用干净语音训练的扩散模型评估语音质量,无需标注。

Non-intrusive Speech Quality Assessment with Diffusion Models Trained on Clean Speech

  • 用纯净语音训练无条件扩散模型,通过噪声过程估计语音似然。
  • 似然值与主观评分相关性最高,优于传统指标。
  • 纯无监督方法,适合无标注语音质量评估场景。

扩散模型在生成高质量自然语音方面表现优异,但其在语音密度估计方面的潜力尚未被充分探索。本文利用仅在纯净语音上训练的无条件扩散模型,对语音质量进行评估。我们发现,可通过确定性去噪过程得到终止高斯分布中的样本似然值来评估语音质量。该方法完全无监督,仅需纯净语音训练,不依赖任何标注。基于清洁语音先验,通过输入与学习到的干净数据分布的关系评估质量。实验表明,所提出的对数似然值与侵入式语音质量指标相关性良好,在听觉实验中与人类评分的相关性达到最佳。

原文摘要 · Abstract (English)

Diffusion models have found great success in generating high quality, natural samples of speech, but their potential for density estimation for speech has so far remained largely unexplored. In this work, we leverage an unconditional diffusion model trained only on clean speech for the assessment of speech quality. We show that the quality of a speech utterance can be assessed by estimating the likelihood of a corresponding sample in the terminating Gaussian distribution, obtained via a deterministic noising process. The resulting method is purely unsupervised, trained only on clean speech, and therefore does not rely on annotations. Our diffusion-based approach leverages clean speech priors to assess quality based on how the input relates to the learned distribution of clean data. Our proposed log-likelihoods show promising results, correlating well with intrusive speech quality metrics and showing the best correlation with human scores in a listening experiment.

语音质量扩散模型无监督密度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。