用旋律引导文本生成音乐,小模型也能超大模型效果。
Melody-Guided Music Generation
- 用对比学习对齐文本与旋律,融合隐式旋律信息
- 仅用1/3参数和1/200数据超越顶尖模型
- 适合音乐创作、影视配乐等需要精准控制的场景
我们提出梅洛迪引导音乐生成(MG2)模型,通过旋律引导文本到音乐的生成。首先,利用新提出的对比语言-音乐预训练方法,将文本与音频波形及其关联旋律对齐,使学习到的文本表示融合了隐含的旋律信息。随后,将检索增强的扩散模块同时基于文本提示和检索到的旋律进行条件化,使生成音乐既反映文本内容,又在显式旋律引导下保持内在和声结构。我们在MusicCaps和MusicBench两个公开数据集上进行了大量实验,结果表明,尽管参数少于1/3、训练数据不足1/200,该模型仍优于当前开源文本到音乐生成模型。此外,我们通过三类用户、五种视角的人工评估,使用新设计问卷探索了其潜在实际应用价值。
原文摘要 · Abstract (English)
We present the Melody-Guided Music Generation (MG2) model, a novel approach using melody to guide the text-to-music generation that, despite a simple method and limited resources, achieves excellent performance. Specifically, we first align the text with audio waveforms and their associated melodies using the newly proposed Contrastive Language-Music Pretraining, enabling the learned text representation fused with implicit melody information. Subsequently, we condition the retrieval-augmented diffusion module on both text prompt and retrieved melody. This allows MG2 to generate music that reflects the content of the given text description, meantime keeping the intrinsic harmony under the guidance of explicit melody information. We conducted extensive experiments on two public datasets: MusicCaps and MusicBench. Surprisingly, the experimental results demonstrate that the proposed MG2 model surpasses current open-source text-to-music generation models, achieving this with fewer than 1/3 of the parameters or less than 1/200 of the training data compared to state-of-the-art counterparts. Furthermore, we conducted comprehensive human evaluations involving three types of users and five perspectives, using newly designed questionnaires to explore the potential real-world applications of MG2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。