arXiv:2604.09371eess.AS2026-04

用语言模型生成离散音符,实现高质量多轨音乐分离

Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models

论文配图:Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models
图 1 · 摘自论文原文
  • 将音乐分离任务转为条件离散令牌生成,通过自回归方式逐步输出音频
  • 在MUSDB18-HQ上语音分离的NISQA得分最高,感知质量接近顶尖判别方法
  • 支持多轨顺序生成,可学习的编码器提升分离效果,适合音乐处理研究者

我们提出一种生成式多轨音乐源分离框架,将任务重定义为条件离散令牌生成。与传统在时域或频域直接估计连续信号的方法不同,本方法结合基于Conformer的条件编码器、双路径神经音频编解码器(HCodec)和仅解码器的语言模型,自回归生成四条目标音轨的音频令牌。生成的令牌通过编解码器解码回波形。在MUSDB18-HQ基准上的评估表明,该生成式方法在感知质量上接近当前最优的判别式方法,同时在人声轨道上取得了最高的NISQA分数。消融实验验证了可学习的Conformer编码器的有效性以及跨轨顺序生成的优势。

原文摘要 · Abstract (English)

We propose a generative framework for multi-track music source separation (MSS) that reformulates the task as conditional discrete token generation. Unlike conventional approaches that directly estimate continuous signals in the time or frequency domain, our method combines a Conformer-based conditional encoder, a dual-path neural audio codec (HCodec), and a decoder-only language model to autoregressively generate audio tokens for four target tracks. The generated tokens are decoded back to waveforms through the codec decoder. Evaluation on the MUSDB18-HQ benchmark shows that our generative approach achieves perceptual quality approaching state-of-the-art discriminative methods, while attaining the highest NISQA score on the vocals track. Ablation studies confirm the effectiveness of the learnable Conformer encoder and the benefit of sequential cross-track generation.

音乐分离生成模型离散建模语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。