arXiv:2511.20224cs.SDcs.AI2025-11

DuoTok让多轨音乐分轨更清晰,兼顾音质与模型预测能力。

DuoTok: Source-Aware Dual-Track Tokenization for Multi-Track Music Language Modeling

  • 分阶段解耦声源,用双码本路由提升分轨信息表达
  • 在0.75 kbps下保持良好音质,同时实现最低cnBPT值
  • 适合需要跨轨结构建模的音乐生成与分析任务

音频分词将连续波形转化为多轨音乐语言模型的离散表示。在双轨建模中,分词需同时满足高保真重建、语言模型强可预测性及跨轨对应关系。本文提出DuoTok,一种源感知双轨分词方法,通过分阶段解耦实现该权衡。首先预训练语义编码器,再以多任务监督正则化,冻结编码器后采用硬双码本路由,并对量化码保持辅助目标。扩散解码器重建高频细节,使分词聚焦于结构化信息用于序列建模。在标准基准上,DuoTok实现最优的可预测性-保真度平衡,在0.75 kbps下达到最低cnBPT值,且重建性能仍具竞争力。在固定双轨语言建模协议下,enBPT亦有提升,表明其增益非仅源于码本大小。受控诊断显示,跨轨干扰下可预测性成本更高,长上下文带来更大收益,说明使用DuoTok训练的模型能利用跨轨结构和非局部历史。

原文摘要 · Abstract (English)

Audio tokenization bridges continuous waveforms and multi-track music language models. In dual-track modeling, tokens should preserve three properties at once: high-fidelity reconstruction, strong predictability under a language model, and cross-track correspondence. We introduce DuoTok, a source-aware dual-track tokenizer that addresses this trade-off through staged disentanglement. DuoTok first pretrains a semantic encoder, then regularizes it with multi-task supervision, freezes the encoder, and applies hard dual-codebook routing while keeping auxiliary objectives on quantized codes. A diffusion decoder reconstructs high-frequency details, allowing tokens to focus on structured information for sequence modeling. On standard benchmarks, DuoTok achieves a favorable predictability-fidelity trade-off, reaching the lowest cnBPT while maintaining competitive reconstruction at 0.75 kbps. Under a held-constant dual-track language modeling protocol, enBPT also improves, indicating gains beyond codebook size effects. Controlled diagnostics show larger predictability costs under cross-track corruption and larger gains from longer context, suggesting that models trained on DuoTok tokens use cross-track structure and non-local history.

音乐生成分词双轨建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。