首个专为音乐源恢复设计的基准数据集,解决真实混音还原难题。
MSRBench: A Benchmarking Dataset for Music Source Restoration
- 构建含原始音源与专业混音的成对数据集
- 引入12类真实世界退化,评估恢复保真度
- 现有模型恢复效果有限,需专用架构
音乐源恢复(MSR)将源分离拓展至包含均衡、压缩、混响等制作效果及真实退化的场景,旨在还原原始未处理音源。现有基准无法衡量恢复保真度:合成数据集使用未处理音源但不真实的混音,真实制作数据集仅提供已处理音源而无干净参考。我们提出首个专为MSR评估设计的基准——MSRBench,包含八个乐器类别下的原始音源-混音对,混音由专业混音师制作。该原始-处理配对可直接评估分离精度与恢复保真度。此外,混音还叠加了12类真实退化,涵盖模拟噪声、声学环境和有损编码。基于U-Net与BSRNN的基线实验分别取得SI-SNR -37.8 dB与-23.4 dB,感知质量(FAD CLAP)约为0.7–0.8,表明仍有巨大提升空间,亟需专用于恢复的模型架构。
原文摘要 · Abstract (English)
Music Source Restoration (MSR) extends source separation to realistic settings where signals undergo production effects (equalization, compression, reverb) and real-world degradations, with the goal of recovering the original unprocessed sources. Existing benchmarks cannot measure restoration fidelity: synthetic datasets use unprocessed stems but unrealistic mixtures, while real production datasets provide only already-processed stems without clean references. We present MSRBench, the first benchmark explicitly designed for MSR evaluation. MSRBench contains raw stem-mixture pairs across eight instrument classes, where mixtures are produced by professional mixing engineers. These raw-processed pairs enable direct evaluation of both separation accuracy and restoration fidelity. Beyond controlled studio conditions, the mixtures are augmented with twelve real-world degradations spanning analog artifacts, acoustic environments, and lossy codecs. Baseline experiments with U-Net and BSRNN achieve SI-SNR of -37.8 dB and -23.4 dB respectively, with perceptual quality (FAD CLAP) around 0.7-0.8, demonstrating substantial room for improvement and the need for restoration-specific architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。