MusicFlow用流匹配实现小模型高效文生音乐,零样本适配补全任务。
MusicFlow: Cascaded Flow Matching for Text Guided Music Generation
- 分层流匹配建模语义与声学特征条件分布
- 模型仅需1/5参数和1/5迭代步数,质量超现有大模型
- 零样本支持音乐补全与续写,适合多任务生成应用
我们提出MusicFlow,一种基于流匹配的级联文本到音乐生成模型。利用自监督表征连接文本描述与音乐音频,构建两个流匹配网络以建模语义与声学特征的条件分布。同时,采用掩码预测作为训练目标,使模型可零样本泛化至音乐补全与续写等任务。在MusicCaps数据集上的实验表明,尽管模型规模比现有方法小2~5倍、迭代步数少5倍,生成音乐仍具备更优质量和文本一致性。模型还能在音乐补全与续写任务中达到竞争力表现。代码与模型将公开。
原文摘要 · Abstract (English)
We introduce MusicFlow, a cascaded text-to-music generation model based on flow matching. Based on self-supervised representations to bridge between text descriptions and music audios, we construct two flow matching networks to model the conditional distribution of semantic and acoustic features. Additionally, we leverage masked prediction as the training objective, enabling the model to generalize to other tasks such as music infilling and continuation in a zero-shot manner. Experiments on MusicCaps reveal that the music generated by MusicFlow exhibits superior quality and text coherence despite being over $2\sim5$ times smaller and requiring $5$ times fewer iterative steps. Simultaneously, the model can perform other music generation tasks and achieves competitive performance in music infilling and continuation. Our code and model will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。