arXiv:2511.19734cs.SD2025-11被引 2

评估神经音频编码器的客观质量指标可靠性

Evaluating Objective Speech Quality Metrics for Neural Audio Codecs

  • 通过MUSHRA测试对比主观评分与客观指标
  • 部分指标与人耳感知高度相关,部分表现不佳
  • 为语音编码器评估提供可选指标建议

神经音频编码器因其在生成建模中的应用而受到关注,可在低比特率下实现高保真音频重建。尽管人类听觉测试仍是评估感知质量的金标准,但其耗时且不实用。本文通过在高保真语音信号上开展MUSHRA听觉测试,分析主观评分与常用客观指标之间的相关性。结果表明,部分指标与人耳感知匹配良好,但另一些指标难以捕捉关键失真。研究为使用神经音频编码器进行语音处理时选择合适评估指标提供了实践指导。

原文摘要 · Abstract (English)

Neural audio codecs have gained recent popularity for their use in generative modeling as they offer high-fidelity audio reconstruction at low bitrates. While human listening studies remain the gold standard for assessing perceptual quality, they are time-consuming and impractical. In this work, we examine the reliability of existing objective quality metrics in assessing the performance of recent neural audio codecs. To this end, we conduct a MUSHRA listening test on high-fidelity speech signals and analyze the correlation between subjective scores and widely used objective metrics. Our results show that, while some metrics align well with human perception, others struggle to capture relevant distortions. Our findings provide practical guidance for selecting appropriate evaluation metrics when using neural audio codecs for speech.

音频编码客观评估语音质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。