arXiv:2605.05554eess.AScs.SD2026-05

提出新音频评估方法,更精准捕捉生成音频质量差异

Optimal Transport Audio Distance with Learned Riemannian Ground Metrics

  • 用可学习的黎曼度量修正距离计算,提升对音频特征敏感度
  • 采用熵正则化Sinkhorn运输,比传统方法提升1.9至3.6倍对噪声的感知能力
  • 能给出每条样本的质量诊断,适合需要可解释性的音频评估场景

在音频生成评估中,弗雷谢音频距离(FAD)是带有结构约束的2-Wasserstein距离,其代价函数为固定嵌入回拉,其不变性集合隐藏严重伪影;耦合方式为高斯拟合,会稀释秩-1污染,相较于离散最优传输表现不佳。本文提出最优传输音频距离(OTAD),分别针对两个核心组件设计专用修正机制:使用残差黎曼度量适配器优化代价函数,采用熵正则化Sinkhorn最优传输实现耦合。在四种轴向协议下对比八种编码器,当ε=0.05时,仅改进耦合的对比显示,Sinkhorn对秩-1敏感度比FAD高出1.9至3.6倍。此外,OTAD在与音频质量主观评分(DCASE 2023 Task 7)的相关性上优于基线指标。作为离散传输方案的内在优势,OTAD可提供每样本诊断,其AUROC ≥ 0.86,而标量或核聚合型指标无法实现此能力。

原文摘要 · Abstract (English)

In audio generation evaluation, Fréchet Audio Distance (FAD) is a 2-Wasserstein distance with structural constraints for both primitives: the cost is a frozen embedding pullback whose invariance set hides severe artifacts, and the coupling is a Gaussian fit that dilutes rank-1 contamination relative to discrete OT. We propose Optimal Transport Audio Distance (OTAD), which corrects each primitive with one dedicated mechanism -- a residual Riemannian ground-metric adapter for the cost and entropic Sinkhorn optimal transport for the coupling. Across eight encoders under a four-axis protocol, coupling-only comparisons at $ε= 0.05$ show that Sinkhorn's rank-1 sensitivity exceeds FAD's by a factor of 1.9 to 3.6. Furthermore, OTAD achieves a higher mean Spearman correlation with audio-quality MOS (DCASE 2023 Task 7) than baseline metrics. As an intrinsic benefit of the discrete transport plan, OTAD yields per-sample diagnostics with AUROC $\ge 0.86$, a capability that scalar- or kernel-aggregated metrics structurally lack.

音频评估最优传输生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。