分阶段修复混音音频中的乐器源信号,突破传统线性混合假设限制。
Multi-Stage Music Source Restoration with BandSplit-RoFormer Separation and HiFi++ GAN
- 先用分频RoFormer分离出8个乐器音轨+辅助音轨
- 通过三阶段课程训练提升分离精度,最终在60.1% SDRi上取得效果
- 采用通用再训练+专用专家的HiFi++ GAN修复波形,适合音乐修复任务
音乐源修复(MSR)旨在从完全混音并经过制作处理的音频中恢复原始未处理的乐器音轨,而制作效果和分发失真违背了常见的线性混合假设。本技术报告介绍了CP-JKU团队在MSR ICASSP Challenge 2025中的系统方案。我们的方法将MSR分解为分离与修复两个阶段:首先,使用单个BandSplit-RoFormer分离器预测8个主音轨加一个辅助音轨,采用三阶段课程学习策略,从4音轨微调(使用LoRA)逐步扩展至8音轨;其次,应用一个通用训练后优化为8个乐器特化专家的HiFi++ GAN波形修复器。实验表明,该方法在测试集上达到60.1%的SDRi指标。
原文摘要 · Abstract (English)
Music Source Restoration (MSR) targets recovery of original, unprocessed instrument stems from fully mixed and mastered audio, where production effects and distribution artifacts violate common linear-mixture assumptions. This technical report presents the CP-JKU team's system for the MSR ICASSP Challenge 2025. Our approach decomposes MSR into separation and restoration. First, a single BandSplit-RoFormer separator predicts eight stems plus an auxiliary other stem, and is trained with a three-stage curriculum that progresses from 4-stem warm-start fine-tuning (with LoRA) to 8-stem extension via head expansion. Second, we apply a HiFi++ GAN waveform restorer trained as a generalist and then specialized into eight instrument-specific experts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。