arXiv:2504.13535cs.SDcs.MM2025-04被引 5

MusFlow用条件流匹配实现图文故事多模态音乐生成,让普通人也能轻松创作。

MusFlow: Multimodal Music Generation via Conditional Flow Matching

  • 通过流匹配将图像、文字等多模态信息对齐到音频嵌入空间,指导音乐生成
  • 在新构建的MMusSet数据集上,跨模态生成音乐质量高,支持单模或融合输入
  • 适合多媒体内容创作者、无专业背景的音乐爱好者使用

音乐生成旨在基于多样条件信息生成符合人类审美的音乐片段。尽管已有模型可依据特定文本描述(如风格、流派、乐器)生成音乐,但普通用户因缺乏专业技能或时间难以写出准确提示,限制了实际应用。为此,本文提出MusFlow,一种基于条件流匹配的多模态音乐生成模型。通过多个MLP将多模态条件信息对齐至音频的CLAP嵌入空间,利用条件流匹配在预训练VAE隐空间中重建压缩的梅尔频谱图。MusFlow可从图像、故事文本和音乐描述生成音乐。为收集训练数据,受多智能体协作启发,我们基于微调的Qwen2-VL构建智能标注流程,建立新数据集MMusSet,每个样本包含图像、故事文本、音乐描述和对应乐曲四元组。实验涵盖图像→音乐、故事→音乐、描述→音乐及多模态生成四类任务。结果表明,无论输入为单模态或多模态,MusFlow均能生成高质量音乐。我们希望推动音乐生成在多媒体领域的应用,使创作更普及。生成样本、代码与数据集见musflow.github.io。

原文摘要 · Abstract (English)

Music generation aims to create music segments that align with human aesthetics based on diverse conditional information. Despite advancements in generating music from specific textual descriptions (e.g., style, genre, instruments), the practical application is still hindered by ordinary users' limited expertise or time to write accurate prompts. To bridge this application gap, this paper introduces MusFlow, a novel multimodal music generation model using Conditional Flow Matching. We employ multiple Multi-Layer Perceptrons (MLPs) to align multimodal conditional information into the audio's CLAP embedding space. Conditional flow matching is trained to reconstruct the compressed Mel-spectrogram in the pretrained VAE latent space guided by aligned feature embedding. MusFlow can generate music from images, story texts, and music captions. To collect data for model training, inspired by multi-agent collaboration, we construct an intelligent data annotation workflow centered around a fine-tuned Qwen2-VL model. Using this workflow, we build a new multimodal music dataset, MMusSet, with each sample containing a quadruple of image, story text, music caption, and music piece. We conduct four sets of experiments: image-to-music, story-to-music, caption-to-music, and multimodal music generation. Experimental results demonstrate that MusFlow can generate high-quality music pieces whether the input conditions are unimodal or multimodal. We hope this work can advance the application of music generation in multimedia field, making music creation more accessible. Our generated samples, code and dataset are available at musflow.github.io.

多模态生成音乐生成流匹配扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。