arXiv:2601.13758cs.SD2026-01AAAI被引 2

改进音频生成的信噪比度量,提升与人耳感知的一致性

GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation Tasks

  • 通过引入相位距离项重构信噪比,解决传统方法忽略相位误差的问题
  • 新度量在神经声码器上表现更优,误差预测更贴近真实感知质量
  • 提出两类新损失函数,适用于相位优化与联合优化,适合语音生成研究者

在音频生成领域,信噪比(SNR)长期作为客观评价指标。然而,近期研究表明,SNR及其变体与人类听觉感知的相关性并不稳定,这促使我们思考:为何SNR无法有效衡量音频质量?如何提升其可靠性?本文识别出相位距离测量不足是关键原因,提出通过引入专门设计的相位距离项来重构SNR,得到改进后的度量GOMPSNR。进一步地,基于该公式推导出两类新型损失函数:幅度引导相位精修和幅度-相位联合优化。通过大量实验验证不同损失组合的最优配置。在先进神经声码器上的实验表明,所提出的GOMPSNR相比传统SNR具有更可靠的误差评估能力;所提损失函数显著提升模型性能,最优组合进一步增强整体生成能力。

原文摘要 · Abstract (English)

In the field of audio generation, signal-to-noise ratio (SNR) has long served as an objective metric for evaluating audio quality. Nevertheless, recent studies have shown that SNR and its variants are not always highly correlated with human perception, prompting us to raise the questions: Why does SNR fail in measuring audio quality? And how to improve its reliability as an objective metric? In this paper, we identify the inadequate measurement of phase distance as a pivotal factor and propose to reformulate SNR with specially designed phase-distance terms, yielding an improved metric named GOMPSNR. We further extend the newly proposed formulation to derive two novel categories of loss function, corresponding to magnitude-guided phase refinement and joint magnitude-phase optimization, respectively. Besides, extensive experiments are conducted for an optimal combination of different loss functions. Experimental results on advanced neural vocoders demonstrate that our proposed GOMPSNR exhibits more reliable error measurement than SNR. Meanwhile, our proposed loss functions yield substantial improvements in model performance, and our wellchosen combination of different loss functions further optimizes the overall model capability.

音频生成信噪比相位优化损失函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。