arXiv:2503.01183eess.AS2025-03被引 74

用简单扩散模型10秒生成4分45秒完整歌曲,支持人声与伴奏一体生成。

DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion

  • 基于潜在扩散模型,端到端生成全曲,无需复杂流水线
  • 4分45秒歌曲生成仅需10秒,保持高音乐性与可懂度
  • 只需歌词和风格提示,适合快速创作与研究复现

近期音乐生成进展引人关注,但现有方法仍存在关键局限:部分模型仅能生成人声或伴奏;少数可生成两者结合的模型依赖复杂的多阶段级联结构和繁琐数据流程,难以扩展;多数系统仅限短段落生成;而主流语言模型方法推理速度慢。为此,我们提出DiffRhythm,首个基于潜在扩散的歌曲生成模型,可在仅10秒内合成长达4分45秒、包含人声与伴奏的完整歌曲,同时保持高音乐性与可懂度。尽管性能卓越,DiffRhythm设计简洁优雅:无需复杂数据预处理,模型结构直接,推理时仅需歌词与风格提示。其非自回归结构确保高速推理,提升可扩展性。我们还公开了完整训练代码及在大规模数据上预训练的模型,以促进可复现性与后续研究。

原文摘要 · Abstract (English)

Recent advancements in music generation have garnered significant attention, yet existing approaches face critical limitations. Some current generative models can only synthesize either the vocal track or the accompaniment track. While some models can generate combined vocal and accompaniment, they typically rely on meticulously designed multi-stage cascading architectures and intricate data pipelines, hindering scalability. Additionally, most systems are restricted to generating short musical segments rather than full-length songs. Furthermore, widely used language model-based methods suffer from slow inference speeds. To address these challenges, we propose DiffRhythm, the first latent diffusion-based song generation model capable of synthesizing complete songs with both vocal and accompaniment for durations of up to 4m45s in only ten seconds, maintaining high musicality and intelligibility. Despite its remarkable capabilities, DiffRhythm is designed to be simple and elegant: it eliminates the need for complex data preparation, employs a straightforward model structure, and requires only lyrics and a style prompt during inference. Additionally, its non-autoregressive structure ensures fast inference speeds. This simplicity guarantees the scalability of DiffRhythm. Moreover, we release the complete training code along with the pre-trained model on large-scale data to promote reproducibility and further research.

音乐生成扩散模型端到端快速生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。