用电影视频提升文本生成音乐的贴合度,让音乐更懂影视情绪。
Video-Guided Text-to-Music Generation Using Public Domain Movie Collections
- 用电影画面条件控制文本到音乐的生成过程。
- 在36.5小时电影片段上训练,显著提升音乐与画面的情绪和类型匹配度。
- 适合影视配乐、跨模态生成方向的研究者和创作者使用。
尽管音乐生成系统已有进展,但在影视制作中的应用仍受限,因其难以捕捉真实拍摄中视觉内容、对白与情感基调等多重因素对配乐的影响。这主要源于缺乏整合多模态信息的完整数据集。为此,我们构建了开放域电影原声库(OSSL),包含约36.5小时公共领域电影片段,配套高质量原声及人工标注的情绪标签。为验证该数据集的有效性,我们提出一种视频适配器,通过引入视频条件增强基于自回归变换器的文本到音乐模型。实验表明,该方法显著提升MusicGen-Medium在分布一致性、配对保真度以及主观情绪与类型契合度上的表现。为促进可复现性与后续研究,我们公开发布数据集、代码与演示。
原文摘要 · Abstract (English)
Despite recent advancements in music generation systems, their application in film production remains limited, as they struggle to capture the nuances of real-world filmmaking, where filmmakers consider multiple factors-such as visual content, dialogue, and emotional tone-when selecting or composing music for a scene. This limitation primarily stems from the absence of comprehensive datasets that integrate these elements. To address this gap, we introduce Open Screen Soundtrack Library (OSSL), a dataset consisting of movie clips from public domain films, totaling approximately 36.5 hours, paired with high-quality soundtracks and human-annotated mood information. To demonstrate the effectiveness of our dataset in improving the performance of pre-trained models on film music generation tasks, we introduce a new video adapter that enhances an autoregressive transformer-based text-to-music model by adding video-based conditioning. Our experimental results demonstrate that our proposed approach effectively enhances MusicGen-Medium in terms of both objective measures of distributional and paired fidelity, and subjective compatibility in mood and genre. To facilitate reproducibility and foster future work, we publicly release the dataset, code, and demo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。