arXiv:2507.06917eess.AS2025-07被引 5

对比客观指标与人耳听感,发现不同乐器分离效果需用不同评估方法。

Musical Source Separation Bake-Off: Comparing Objective Metrics with Human Perception

  • 在MUSDB18数据集上测试7组听众,每首歌收集约30次评分。
  • SI-SAR对鼓和贝斯的分离质量预测能力优于传统指标SDR。
  • CLAP-LAION-music的FAD值对鼓/贝斯表现良好,但对人声无效。

音乐源分离旨在从混音中提取独立声源(如人声、鼓、吉他)。然而,评估分离音频质量仍具挑战性,因常用指标如信源失真比(SDR)常与人耳感知不一致。本研究在MUSDB18测试集上开展大规模听感评估,从七个不同听众群体收集每首歌约30次评分。比较了多种能量比指标(包括BSSEval v4、SI-SDR变体)及基于嵌入的度量(使用CLAP-LAION-music、EnCodec、VGGish、Wave2Vec2、HuBERT计算的弗雷切特音频距离,即FAD)。尽管SDR仍是人声分离的最佳指标,但尺度不变信号-干扰比(SI-SAR)对鼓和贝斯茎干的听感预测更优。基于CLAP-LAION-music的FAD在鼓和贝斯上的肯德尔等级相关系数分别为0.25和0.19,表现可媲美甚至超越能量型指标;然而,所有嵌入式指标均未与人声听感呈现正相关。结果表明,需针对不同声源制定特定评估策略,且单一指标无法全面反映感知质量。研究公开原始听感评分以支持复现与后续研究。

原文摘要 · Abstract (English)

Music source separation aims to extract individual sound sources (e.g., vocals, drums, guitar) from a mixed music recording. However, evaluating the quality of separated audio remains challenging, as commonly used metrics like the source-to-distortion ratio (SDR) do not always align with human perception. In this study, we conducted a large-scale listener evaluation on the MUSDB18 test set, collecting approximately 30 ratings per track from seven distinct listener groups. We compared several objective energy-ratio metrics, including legacy measures (BSSEval v4, SI-SDR variants), and embedding-based alternatives (Frechet Audio Distance using CLAP-LAION-music, EnCodec, VGGish, Wave2Vec2, and HuBERT). While SDR remains the best-performing metric for vocal estimates, our results show that the scale-invariant signal-to-artifacts ratio (SI-SAR) better predicts listener ratings for drums and bass stems. Frechet Audio Distance (FAD) computed with the CLAP-LAION-music embedding also performs competitively--achieving Kendall's tau values of 0.25 for drums and 0.19 for bass--matching or surpassing energy-based metrics for those stems. However, none of the embedding-based metrics, including CLAP, correlate positively with human perception for vocal estimates. These findings highlight the need for stem-specific evaluation strategies and suggest that no single metric reliably reflects perceptual quality across all source types. We release our raw listener ratings to support reproducibility and further research.

音乐分离评估指标人耳感知听感评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。