arXiv:2606.24127eess.AScs.AI2026-06中稿 · Interspeech 2026

分两阶段生成与重建,提升音乐源分离的音质与语义一致性。

DTT-BSR+: A Generative-Regression Cascade for Music Source Restoration

论文配图:DTT-BSR+: A Generative-Regression Cascade for Music Source Restoration
图 1 · 摘自论文原文
  • 先生成符合干净源分布的音轨,再用时域与多分辨率频谱损失优化信号
  • 所有音轨的多梅尔信噪比均高于单阶段方法,五类音轨超越顶尖模型X-LANCE
  • 揭示了重建精度与语义匹配间的内在权衡,适合音频修复与高质量音乐生成研究者

音乐源恢复(MSR)需同时处理源分离与非线性制作效果的逆向问题。现有方法难以在精确重建目标信号的同时保持语义一致性。为此,我们提出DTT-BSR+,一种两阶段级联式MSR系统,将分布拟合与信号重建分离处理。第一阶段采用生成式DTT-BSR分离器,输出符合干净源先验的音轨;第二阶段通过改进的Demucs网络,结合时域与多分辨率频谱损失增强第一阶段结果。DTT-BSR+在所有音轨上均实现优于单阶段DTT-BSR的多梅尔信噪比(MMSNR),并在五类音轨上超越当前最优的X-LANCE MSR系统。通过弗雷谢音频距离(FAD)分解,我们揭示了音轨间信号重建精度与语义分布拟合之间的隐含权衡。

原文摘要 · Abstract (English)

Music source restoration (MSR) requires jointly addressing source unmixing and the inversion of non-linear production effects. Current methods struggle to achieve accurate target signal reconstruction while maintaining semantic consistency. To address this limitation, we propose DTT-BSR+, a two-stage cascade MSR system that decouples distribution fitting from signal reconstruction into separate stages. A generative DTT-BSR separator in the first stage produces stems matching the prior of clean sources, and a modified Demucs network in the second stage enhances the first stage output using time-domain and multi-resolution spectral losses. DTT-BSR+ improves multi-mel signal-to-noise ratio (MMSNR) over the single-stage DTT-BSR across all stems, and surpasses the state-of-the-art X-LANCE MSR system on five stems. We also reveal through Fréchet Audio Distance (FAD) decomposition an implicit trade-off between signal reconstruction accuracy and semantic distribution fitting across stems.

音乐修复源分离生成模型信号重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。