纯扩散模型一键生成5分钟高保真多语言歌曲,支持音轨分离。
WanSong v1.0 Technical Report

- 纯扩散架构直接生成长音频,无需分步处理。
- 单次运行输出人声与伴奏双轨,最长可达5分钟。
- 支持快速推理与下游编辑,适合商业音乐生成场景。
音乐生成基础模型近期受到业界广泛关注。然而,在实现高效生成、高保真长时音频并保持可控性方面仍面临挑战。为解决上述问题,我们提出WanSong,一种简单但强大的长时、商用级歌曲生成方法。不同于自回归(AR)或级联多阶段流程(如先自回归再扩散),WanSong是一种纯扩散模型,可直接生成高达5分钟的高保真多语言歌曲,并在单次运行中输出人声与背景音乐双轨。此外,我们的扩散框架通过步骤蒸馏实现更快推理,并为微调与定制化提供高效路径,支持下游编辑任务。
原文摘要 · Abstract (English)
Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present \textbf{WanSong}, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (\eg, AR followed by diffusion), \textbf{WanSong} is a pure diffusion-based model that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems (vocals and background music) in a single run. In addition, our diffusion framework enables faster inference through step-distillation, and offers an efficient pathway for fine-tuning and customization to support downstream editing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。