用图文桥提升跨模态音乐生成质量与可控性
Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation

- 通过视觉转文本+双轨检索构建显式图文桥梁
- 在多任务测试中显著提升音乐质量与对齐度
- 适合需要精细控制的音乐创作与多媒体应用
多模态音乐生成旨在从文本、视频、图像等多元输入生成音乐。现有方法依赖共享嵌入空间进行模态融合,但在音乐生成中面临数据稀缺、跨模态对齐弱和可控性差的问题。本文提出视觉音乐桥(VMB)方法:先用多模态音乐描述模型将视觉输入转为详细文本描述作为文本桥;再通过双轨音乐检索模块结合广义与精准检索策略提供音乐桥并实现用户控制;最后基于显式条件生成框架生成音乐。在视频到音乐、图像到音乐、文本到音乐及可控音乐生成任务上验证,结果表明VMB显著提升音乐质量、模态对齐与定制化一致性。该方法为可解释且富有表现力的多模态音乐生成树立新标准,适用于多种多媒体场景。演示与代码见https://github.com/wbs2788/VMB。
原文摘要 · Abstract (English)
Multimodal music generation aims to produce music from diverse input modalities, including text, videos, and images. Existing methods use a common embedding space for multimodal fusion. Despite their effectiveness in other modalities, their application in multimodal music generation faces challenges of data scarcity, weak cross-modal alignment, and limited controllability. This paper addresses these issues by using explicit bridges of text and music for multimodal alignment. We introduce a novel method named Visuals Music Bridge (VMB). Specifically, a Multimodal Music Description Model converts visual inputs into detailed textual descriptions to provide the text bridge; a Dual-track Music Retrieval module that combines broad and targeted retrieval strategies to provide the music bridge and enable user control. Finally, we design an Explicitly Conditioned Music Generation framework to generate music based on the two bridges. We conduct experiments on video-to-music, image-to-music, text-to-music, and controllable music generation tasks, along with experiments on controllability. The results demonstrate that VMB significantly enhances music quality, modality, and customization alignment compared to previous methods. VMB sets a new standard for interpretable and expressive multimodal music generation with applications in various multimedia fields. Demos and code are available at https://github.com/wbs2788/VMB.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。