用MERT音频嵌入改进歌声分离评估,更贴近人耳听感。
Embedding-Based Intrusive Evaluation Metrics for Musical Source Separation Using MERT Representations
- 基于MERT音频嵌入计算误差和距离度量
- 在两个数据集上与听感评分相关性更高
- 适合追求真实听感评估的研究者
音乐源分离(MSS)传统上依赖盲源分离评估(BSS-Eval)指标。然而,近期研究表明,这些指标与听觉测试中的人耳感知质量评分相关性较低,而听觉测试被视为金标准。为此,有研究提出利用大规模自监督音频模型(如MERT)的潜在表示,构建基于嵌入的侵入式评估指标。本文分析了两种基于嵌入的侵入式指标:基于均方误差(MSE)和在MERT嵌入上计算的侵入式弗雷歇音频距离(FAD)与感知音频质量评分的相关性。在两个独立数据集上的实验表明,这些指标在所有分析的音轨类型和模型类型下,均比传统BSS-Eval指标与听感评分具有更强的相关性。
原文摘要 · Abstract (English)
Evaluation of musical source separation (MSS) has traditionally relied on Blind Source Separation Evaluation (BSS-Eval) metrics. However, recent work suggests that BSS-Eval metrics exhibit low correlation between metrics and perceptual audio quality ratings from a listening test, which is considered the gold standard evaluation method. As an alternative approach in singing voice separation, embedding-based intrusive metrics that leverage latent representations from large self-supervised audio models such as Music undERstanding with large-scale self-supervised Training (MERT) embeddings have been introduced. In this work, we analyze the correlation of perceptual audio quality ratings with two intrusive embedding-based metrics: a mean squared error (MSE) and an intrusive variant of the Fréchet Audio Distance (FAD) calculated on MERT embeddings. Experiments on two independent datasets show that these metrics correlate more strongly with perceptual audio quality ratings than traditional BSS-Eval metrics across all analyzed stem and model types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。