arXiv:2507.11427eess.AS2025-07中稿 · presentation at th…被引 6

提出更可靠的评估生成式声乐分离模型的方法,发现传统指标不适用。

Towards Reliable Objective Evaluation Metrics for Generative Singing Voice Separation Models

  • 用听觉测试验证多种客观指标有效性
  • 生成模型用MERT-L12嵌入的MSE相关性最高
  • 建议用嵌入空间误差替代传统度量

传统盲源分离评估(BSS-Eval)指标原本针对基于时频掩码的线性分离模型设计,但近期生成模型引入了分离信号与参考信号间的非线性关系,使这些指标在客观评估中可靠性下降。为此,我们开展降级类别评分听觉测试,分析获得的降级平均意见分(DMOS)与多种客观音频质量指标的相关性,评估了三种先进判别式模型和两种新型生成式模型。对判别式模型,基于嵌入的侵入式指标(如音乐2潜在表示上的MSE)比传统指标(如BSS-Eval)相关性更高;对生成式模型,多分辨率STFT损失和MERT-L12嵌入上的MSE相关性最强,后者在两类模型间表现最均衡。结果表明BSS-Eval指标难以有效评估生成式声乐分离模型,强调需谨慎选择并验证替代评估指标。

原文摘要 · Abstract (English)

Traditional Blind Source Separation Evaluation (BSS-Eval) metrics were originally designed to evaluate linear audio source separation models based on methods such as time-frequency masking. However, recent generative models may introduce nonlinear relationships between the separated and reference signals, limiting the reliability of these metrics for objective evaluation. To address this issue, we conduct a Degradation Category Rating listening test and analyze correlations between the obtained degradation mean opinion scores (DMOS) and a set of objective audio quality metrics for the task of singing voice separation. We evaluate three state-of-the-art discriminative models and two new competitive generative models. For both discriminative and generative models, intrusive embedding-based metrics show higher correlations with DMOS than conventional intrusive metrics such as BSS-Eval. For discriminative models, the highest correlation is achieved by the MSE computed on Music2Latent embeddings. When it comes to the evaluation of generative models, the strongest correlations are evident for the multi-resolution STFT loss and the MSE calculated on MERT-L12 embeddings, with the latter also providing the most balanced correlation across both model types. Our results highlight the limitations of BSS-Eval metrics for evaluating generative singing voice separation models and emphasize the need for careful selection and validation of alternative evaluation metrics for the task of singing voice separation.

声乐分离生成模型评估指标音频质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。