arXiv:2502.15602cs.SDcs.AI2025-02被引 39

KAD替代传统音频评估指标,更快更准且省算力。

KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation

  • 基于最大均值差异,无需假设分布形态。
  • 小样本下收敛快,计算成本低,支持GPU加速。
  • 更贴近人耳判断,适合评估生成音频质量。

尽管广泛用于评估生成音频信号,弗雷谢音频距离(FAD)存在严重局限:依赖高斯假设、对样本量敏感、计算复杂度高。为此,我们提出核音频距离(KAD),一种基于最大均值差异(MMD)的新度量,具有分布无关、无偏、计算高效的特点。通过分析与实证验证,KAD表现出显著优势:(1)在较小样本量下更快收敛,支持有限数据下的可靠评估;(2)计算开销更低,具备可扩展的GPU加速能力;(3)与人类听觉感知判断更强相关。通过利用先进嵌入和特征核,KAD能捕捉真实与生成音频间的细微差异。KAD已开源至kadtk工具包,为生成音频模型提供高效、可靠且感知一致的评估基准。

原文摘要 · Abstract (English)

Although being widely adopted for evaluating generated audio signals, the Fréchet Audio Distance (FAD) suffers from significant limitations, including reliance on Gaussian assumptions, sensitivity to sample size, and high computational complexity. As an alternative, we introduce the Kernel Audio Distance (KAD), a novel, distribution-free, unbiased, and computationally efficient metric based on Maximum Mean Discrepancy (MMD). Through analysis and empirical validation, we demonstrate KAD's advantages: (1) faster convergence with smaller sample sizes, enabling reliable evaluation with limited data; (2) lower computational cost, with scalable GPU acceleration; and (3) stronger alignment with human perceptual judgments. By leveraging advanced embeddings and characteristic kernels, KAD captures nuanced differences between real and generated audio. Open-sourced in the kadtk toolkit, KAD provides an efficient, reliable, and perceptually aligned benchmark for evaluating generative audio models.

音频生成评估指标KADMMD

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。