无需参考音频和背景集,用流匹配逆推评估音乐质量
InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion

- 用简单欧拉积分实现无条件流匹配逆推
- 在多个生成模型上与人耳判断高度相关(相关性>0.9)
- 适合评估音乐生成质量,无需复杂数据准备
现有无参考音乐质量评估方法虽免去了成对噪声-纯净音频的需求,但仍依赖背景集来计算纯净音频的统计特征。本文提出新方法,完全消除对背景集和参考音频的依赖,仅使用预训练的流匹配骨干网络即可实现质量评估。通过简单的欧拉积分进行无条件流匹配逆推,可有效检测各类人为失真,并准确排序音乐生成模型的表现。我们提出InvFlowFD,通过流逆推生成样本并比较其与先验分布的差异。在定量实验和详尽的人类听觉研究中验证,InvFlowFD与人类对声音失真的感知及生成模型质量的评价高度一致,且比现有方法更灵活、限制更少。
原文摘要 · Abstract (English)
Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely on a background set, which is used to compute aggregated statistics of clean audio samples. In this work, we propose a novel approach that eliminates this requirement, achieving background-set-free and reference-free quality estimation using only a pre-trained Flow Matching backbone. We demonstrate that unconditional Flow Matching inversion via simple Euler integration is sufficient to detect various artificial distortions and accurately rank music generation models against human perceptual judgments. We introduce InvFlowFD, which performs flow inversion and compares a group of inverted samples to the prior distribution. We evaluate our method against prior work, quantitatively and with a thorough human study. Results suggest that InvFlowFD is highly correlated with human perception of sound distortions, as well as generative models' quality, while being more flexible and less restrictive than existing metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。