提出新评估方法,让音乐生成模型评价更贴近人类喜好。
Aligning Text-to-Music Evaluation with Human Preferences
- 用自监督音频嵌入构建新指标MAD,替代传统FAD
- MAD在合成数据和真人偏好上相关性达0.84,远超FAD的0.49
- 适合关注音乐生成质量评估的研究者与开发者
尽管生成式声学文本到音乐(TTM)建模取得显著进展,但其评估仍滞后,主要依赖广泛使用的弗雷谢音频距离(FAD)。本文通过(1)设计四种合成元评估以衡量对特定音乐特性的敏感性,及(2)收集并评估首个开源的人类偏好数据集MusicPrefs,系统研究了基于参考的差异度量设计空间。发现标准FAD设置在合成与人类偏好数据上均不一致,且几乎所有现有指标均无法有效捕捉理想音乐特性,与人类感知弱相关。为此,提出基于自监督音频嵌入模型表示的新指标MAUVE Audio Divergence(MAD)。实验表明,MAD能有效捕捉多样音乐特性(平均秩相关性0.84),显著优于FAD的0.49;在MusicPrefs上的相关性也从0.14提升至0.62。
原文摘要 · Abstract (English)
Despite significant recent advances in generative acoustic text-to-music (TTM) modeling, robust evaluation of these models lags behind, relying in particular on the popular Fréchet Audio Distance (FAD). In this work, we rigorously study the design space of reference-based divergence metrics for evaluating TTM models through (1) designing four synthetic meta-evaluations to measure sensitivity to particular musical desiderata, and (2) collecting and evaluating on MusicPrefs, the first open-source dataset of human preferences for TTM systems. We find that not only is the standard FAD setup inconsistent on both synthetic and human preference data, but that nearly all existing metrics fail to effectively capture desiderata, and are only weakly correlated with human perception. We propose a new metric, the MAUVE Audio Divergence (MAD), computed on representations from a self-supervised audio embedding model. We find that this metric effectively captures diverse musical desiderata (average rank correlation 0.84 for MAD vs. 0.49 for FAD and also correlates more strongly with MusicPrefs (0.62 vs. 0.14).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。