arXiv:2509.18823eess.AS2025-09中稿 · 51st IEEE Internat…被引 1

用神经音频编码器评估生成音频质量,效果优于传统方法。

Towards Evaluating Generative Audio: Insights from Neural Audio Codec Embedding Distances

  • 用高保真编码器提取音频嵌入,用于感知质量评估。
  • FAD指标比MMD更贴近人耳判断,且与主观评分相关性更强。
  • 无需标注数据即可零样本评估,适合工业级音频质量检测。

神经音频编码器(NACs)通过学习紧凑的音频表示实现低比特率压缩,也可作为感知质量评估的特征。本文提出DACe,是基于多样化真实与合成音调数据、采用均衡采样训练的描述音频编码器(DAC)增强版。系统比较了在语音、音乐及混合内容上的MUSHRA测试中,弗雷歇音频距离(FAD)与最大均值差异(MMD)的表现。结果表明,FAD始终优于MMD,且来自更高保真度NAC(如DACe)的嵌入与人类判断的相关性更强。尽管CLAP LAION Music(CLAP-M)和OpenL3 Mel128(OpenL3-128M)嵌入相关性更高,但NAC嵌入提供了一种仅需未编码音频即可训练的实用零样本音频质量评估方法。这些结果证明了NAC在压缩与感知评估中的双重价值。

原文摘要 · Abstract (English)

Neural audio codecs (NACs) achieve low-bitrate compression by learning compact audio representations, which can also serve as features for perceptual quality evaluation. We introduce DACe, an enhanced, higher-fidelity version of the Descript Audio Codec (DAC), trained on diverse real and synthetic tonal data with balanced sampling. We systematically compare Fréchet Audio Distance (FAD) and Maximum Mean Discrepancy (MMD) on MUSHRA tests across speech, music, and mixed content. FAD consistently outperforms MMD, and embeddings from higher-fidelity NACs (such as DACe) show stronger correlations with human judgments. While CLAP LAION Music (CLAP-M) and OpenL3 Mel128 (OpenL3-128M) embeddings achieve higher correlations, NAC embeddings provide a practical zero-shot approach to audio quality assessment, requiring only unencoded audio for training. These results demonstrate the dual utility of NACs for compression and perceptually informed audio evaluation.

音频评估神经编码器感知质量零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。