arXiv:2503.08638eess.AScs.AI2025-03被引 101

开源模型YuE实现长达5分钟的歌词到歌曲生成,保持歌词对齐与音乐连贯性。

YuE: Scaling Open Foundation Models for Long-Form Music Generation

  • 采用分轨预测与渐进式结构条件化,解决长序列音乐生成中的信号干扰问题。
  • 在五分鐘音樂生成中實現高對齊度、連貫結構與自然人聲旋律,媲美專有系統。
  • 支持風格轉換與雙向生成,適合音樂創作、跨語言創作與音樂理解研究者。

我們針對長時序音樂生成——特別是具有挑戰性的「歌詞到歌曲」問題——提出基於LLaMA2架構的開源基礎模型家族YuE。YuE訓練規模達數萬億標記,可生成最長五分鐘的音樂,同時維持歌詞對齊、連貫的音樂結構與吸引人的演唱旋律及合適伴奏。其技術核心包括:(1)分軌解耦的下一個詞元預測以克服密集混合信號;(2)結構性逐步條件化以實現長上下文歌詞對齊;(3)多任務、多階段預訓練策略以促進收斂與泛化。此外,我們重新設計了音樂生成的上下文學習技術,實現多樣風格轉換(如將日本城市流行轉為英文說唱,保留原伴奏)與雙向生成。大量實驗表明,YuE在音樂性與人聲靈活性上達到甚至超越部分專有系統。微調後,模型可實現額外控制並增強對尾部語言的支持。進一步實驗顯示,YuE學得的表示在音樂理解任務中表現優異,在MARBLE基準上達到或超過現有最佳方法。

原文摘要 · Abstract (English)

We tackle the task of long-form music generation--particularly the challenging \textbf{lyrics-to-song} problem--by introducing YuE, a family of open foundation models based on the LLaMA2 architecture. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while maintaining lyrical alignment, coherent musical structure, and engaging vocal melodies with appropriate accompaniment. It achieves this through (1) track-decoupled next-token prediction to overcome dense mixture signals, (2) structural progressive conditioning for long-context lyrical alignment, and (3) a multitask, multiphase pre-training recipe to converge and generalize. In addition, we redesign the in-context learning technique for music generation, enabling versatile style transfer (e.g., converting Japanese city pop into an English rap while preserving the original accompaniment) and bidirectional generation. Through extensive evaluation, we demonstrate that YuE matches or even surpasses some of the proprietary systems in musicality and vocal agility. In addition, fine-tuning YuE enables additional controls and enhanced support for tail languages. Furthermore, beyond generation, we show that YuE's learned representations can perform well on music understanding tasks, where the results of YuE match or exceed state-of-the-art methods on the MARBLE benchmark. Keywords: lyrics2song, song generation, long-form, foundation model, music generation

音樂生成長序列開源模型歌词到歌曲

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。