arXiv:2502.10154cs.SDcs.AI2025-02被引 2

让音乐精准匹配视频情绪与镜头切换,自动生成同步感更强的配乐。

Video Soundtrack Generation by Aligning Emotions and Temporal Boundaries

  • 分两阶段生成:先识情绪,再用情绪与时间信号引导作曲。
  • 在多个数据集上超越现有模型,主观评价更受欢迎。
  • 创新设计时间偏移机制,提前对齐音乐与画面转折点。

为视频配乐仍是多媒体内容创作者面临的高成本、耗时挑战。我们提出EMSYNC,一种基于视频的符号化音乐自动生成方法,能生成与视频情绪内容和时间边界精确对齐的音乐。该方法采用两阶段框架:首先使用预训练视频情绪分类器提取情绪特征,再由条件音乐生成器根据情绪与时间线索生成MIDI序列。我们引入边界偏移(boundary offsets)这一新颖的时间条件机制,使模型能够提前感知即将发生的视频场景切换,并将生成的和弦与之对齐。此外,我们提出一种映射方案,将视频情绪分类器的离散类别输出与情绪条件化MIDI生成器所需的连续唤醒-效价输入进行衔接,实现不同表示间的情绪信息无缝融合。我们的方法在多个视频数据集上的客观与主观评估中均优于当前最优模型,证明其在情感与时间双重对齐方面的有效性。演示及生成样本见https://serkansulun.com/emsync。

原文摘要 · Abstract (English)

Providing soundtracks for videos remains a costly and time-consuming challenge for multimedia content creators. We introduce EMSYNC, an automatic video-based symbolic music generator that creates music aligned with a video's emotional content and temporal boundaries. It follows a two-stage framework, where a pretrained video emotion classifier extracts emotional features, and a conditional music generator produces MIDI sequences guided by both emotional and temporal cues. We introduce boundary offsets, a novel temporal conditioning mechanism that enables the model to anticipate upcoming video scene cuts and align generated musical chords with them. We also propose a mapping scheme that bridges the discrete categorical outputs of the video emotion classifier with the continuous valence-arousal inputs required by the emotion-conditioned MIDI generator, enabling seamless integration of emotion information across different representations. Our method outperforms state-of-the-art models in objective and subjective evaluations across different video datasets, demonstrating its effectiveness in generating music aligned to video both emotionally and temporally. Our demo and output samples are available at https://serkansulun.com/emsync.

视频配乐情绪对齐时间对齐MIDI生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。