arXiv:2503.03304eess.AS2025-03被引 2

提出新指标LQR,用潜在表示距离衡量语音质量,效果媲美甚至超越传统方法。

On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs

  • 引入LQR指标,量化神经编码器潜在表示与理想模型的距离。
  • LQR与主观语音质量相关性高达0.9(皮尔逊相关系数)。
  • 非侵入式设计,适合实时语音质量评估,尤其适合无参考场景。

近年来,神经音频编码器受到广泛关注。其低码率表现源于学习到能捕捉信号(如语音)特性的抽象表示。本文研究神经编码器学习的潜在表示与语音质量之间的关系。为此,提出潜表示到量化误差比(LQR)度量,用于量化给定语音信号与理想神经编码器语音模型之间的偏离程度。在两个主观语音质量数据集上,将LQR与侵入式度量及数据驱动的监督方法进行比较。结果表明,所提LQR与主观语音质量具有强相关性(最高达0.9皮尔逊相关系数)。尽管是非侵入式度量,其性能仍可媲美甚至优于其他预训练和侵入式度量。这些结果表明,LQR是构建更复杂语音质量度量的有前景基础。

原文摘要 · Abstract (English)

Neural audio signal codecs have attracted significant attention in recent years. In essence, the impressive low bitrate achieved by such encoders is enabled by learning an abstract representation that captures the properties of encoded signals, e.g., speech. In this work, we investigate the relation between the latent representation of the input signal learned by a neural codec and the quality of speech signals. To do so, we introduce Latent-representation-to-Quantization error Ratio (LQR) measures, which quantify the distance from the idealized neural codec's speech signal model for a given speech signal. We compare the proposed metrics to intrusive measures as well as data-driven supervised methods using two subjective speech quality datasets. This analysis shows that the proposed LQR correlates strongly (up to 0.9 Pearson's correlation) with the subjective quality of speech. Despite being a non-intrusive metric, this yields a competitive performance with, or even better than, other pre-trained and intrusive measures. These results show that LQR is a promising basis for more sophisticated speech quality measures.

语音质量神经编码器无参考评估潜在表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。