arXiv:2608.11576cs.SDcs.CV2026-08被引 1

用公开电影数据集提升视频配乐质量,让音乐更贴合对话节奏。

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

论文配图:Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections
图 1 · 摘自论文原文
  • 用公开电影片段构建可复现的音视频数据集,避免链接失效。
  • 在视频注意力中加入对话时间轴,使配乐随对话帧同步变化。
  • 适合影视配乐、跨模态生成研究者,数据已开源可用。

视频到音乐生成因能传达视觉媒体的情感而受到关注,但当前研究受限于可复现性:模型常基于可通过YouTube链接获取的数据集训练,这些链接可能失效,且原始数据难获取。为此,我们提出Open Screen Soundtrack Library v2(OSSL-v2),一个自托管的语料库,包含34,343个视频片段,总计246.4小时,均来自公共领域电影。与爬取数据集不同,OSSL-v2具有可复现性和版权合规性,同时规模足够支持功能型视频到音乐模型训练。我们利用该数据集研究对话作为视频配乐的条件信号,动机在于电影音乐与对白存在紧密的时间耦合。具体地,在现有模型的视频交叉注意力中引入时间轴,并通过对话轨道逐帧调节。在公共领域及商业电影上评估,本方法优于当前最优基线。数据集已发布于https://huggingface.co/datasets/McAuley-Lab/OSSL-v2。

原文摘要 · Abstract (English)

Video-to-music generation has drawn growing interest for its role in conveying the emotion of visual media, including film. Progress in the field, however, is hampered by a reproducibility gap: models are often trained on crawled corpora referenced through YouTube URLs that may be deleted, with the underlying data often difficult and time-consuming to retrieve. To address this, we introduce the Open Screen Soundtrack Library version 2 (OSSL-v2), a self-hosted corpus of 34,343 video clips totaling 246.4 hours, sourced from public-domain films. Unlike crawled corpora, OSSL-v2 is reproducible (i.e., not subject to link rot) and copyright-conscious, yet still large enough to train functional video-to-music models. We then use this film-domain corpus to study dialogue as a conditioning signal for video-to-music generation, motivated by the close temporal coupling between film music and on-screen speech. Specifically, we augment existing models' video cross-attention with a time axis and modulate it frame-by-frame with the dialogue track. Evaluated on both public-domain and commercial films, our approach shows improvement over the state-of-the-art baselines. The dataset is available at https://huggingface.co/datasets/McAuley-Lab/OSSL-v2.

视频配乐跨模态生成数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。