arXiv:2602.19825eess.AScs.SD2026-02中稿 · ICASSP 2026被引 1

用混合GAN模型修复混音中的音乐源信号,效果出色且参数少。

DTT-BSR: GAN-based DTTNet with RoPE Transformer Enhancement for Music Source Restoration

  • 结合旋转位置编码的Transformer与双路径频带循环网络,兼顾时序与频谱建模。
  • 在ICASSP 2026挑战赛中客观评分第3、主观评分第4,仅710万参数。
  • 适合需要轻量高保真音频修复的研究与应用开发者。

音乐源恢复(MSR)旨在从混音和母带化录音中还原未处理的音轨。难点在于既要分离重叠声源,又要重建因压缩、混响等制作效果退化的信号。为此,我们提出DTT-BSR,一种融合旋转位置编码(RoPE)Transformer用于长期时序建模,以及双路径频带分割循环神经网络(RNN)实现多分辨率频谱处理的混合生成对抗网络(GAN)。该模型在ICASSP 2026 MSR挑战赛中取得客观评价第3名、主观评价第4名的成绩,展现出卓越的生成保真度与语义一致性,模型规模仅为710万参数。

原文摘要 · Abstract (English)

Music source restoration (MSR) aims to recover unprocessed stems from mixed and mastered recordings. The challenge lies in both separating overlapping sources and reconstructing signals degraded by production effects such as compression and reverberation. We therefore propose DTT-BSR, a hybrid generative adversarial network (GAN) combining rotary positional embeddings (RoPE) transformer for long-term temporal modeling with dual-path band-split recurrent neural network (RNN) for multi-resolution spectral processing. Our model achieved 3rd place on the objective leaderboard and 4th place on the subjective leaderboard on the ICASSP 2026 MSR Challenge, demonstrating exceptional generation fidelity and semantic alignment with a compact size of 7.1M parameters.

音乐修复生成模型GANTransformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。