用无配对音视频数据生成适配视频的背景音乐。
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
- 用大模型理解视频并转为音乐标签,再通过扩散模型生成音乐。
- 支持控制乐器、风格、节奏等细节,生成音乐贴合视频内容。
- 无需配对数据,开源可试用,适合内容创作者快速配乐。
我们提出 SONIQUE,一个基于无配对音视频数据生成适配视频背景音乐的模型。与依赖成对音视频数据的传统方法不同,SONIQUE利用免费版权音乐和独立视频源。通过大语言模型(LLMs)理解视频内容,将视觉描述转化为音乐标签,并结合基于U-Net的条件扩散模型,实现可定制的音乐生成。用户可控制乐器、音乐风格、节拍和旋律等要素,确保输出符合创作意图。模型已开源,提供在线演示。
原文摘要 · Abstract (English)
We present SONIQUE, a model for generating background music tailored to video content. Unlike traditional video-to-music generation approaches, which rely heavily on paired audio-visual datasets, SONIQUE leverages unpaired data, combining royalty-free music and independent video sources. By utilizing large language models (LLMs) for video understanding and converting visual descriptions into musical tags, alongside a U-Net-based conditional diffusion model, SONIQUE enables customizable music generation. Users can control specific aspects of the music, such as instruments, genres, tempo, and melodies, ensuring the generated output fits their creative vision. SONIQUE is open-source, with a demo available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。