arXiv:2602.07063cs.LGcs.AI2026-02

根据视频自动生成情感与节奏同步的音乐,无需作曲或买版权。

Video-based Music Generation

  • 用视频情绪分类器+连续情绪条件生成音乐,提升情感匹配度。
  • 在Ekman-6和MovieNet上达当前最优,用户评测胜过现有方法。
  • 适合内容创作者快速配乐,尤其需精准情绪表达的视频场景。

随着网络视频内容激增,寻找合适配乐仍是一大挑战。本论文提出EMSYNC(EMotion and SYNChronization),一种快速、免费、全自动的视频配乐方案,可生成与输入视频情感和节奏同步的音乐,帮助内容创作者无需作曲或授权即可提升作品质量。模型核心是一个新型视频情绪分类器,通过冻结预训练深度神经网络并仅训练融合层,降低计算开销同时提升准确率。我们在Ekman-6和MovieNet数据集上取得当前最优表现,验证了方法的泛化能力。另一关键贡献是构建了一个大规模情绪标注的MIDI数据集,用于情感音乐生成。我们提出首个基于连续情绪值而非离散类别的MIDI生成模型,实现更细腻的情感响应。为增强时间同步,引入“边界偏移编码”机制,使和弦与场景切换对齐。结合情绪分类、连续情绪生成与时间边界建模,EMSYNC成为全自动化视频音乐生成系统。用户研究显示,其在音乐丰富性、情感契合度、时间同步性及整体偏好上均优于现有方法,确立视频配乐新基准。

原文摘要 · Abstract (English)

As the volume of video content on the internet grows rapidly, finding a suitable soundtrack remains a significant challenge. This thesis presents EMSYNC (EMotion and SYNChronization), a fast, free, and automatic solution that generates music tailored to the input video, enabling content creators to enhance their productions without composing or licensing music. Our model creates music that is emotionally and rhythmically synchronized with the video. A core component of EMSYNC is a novel video emotion classifier. By leveraging pretrained deep neural networks for feature extraction and keeping them frozen while training only fusion layers, we reduce computational complexity while improving accuracy. We show the generalization abilities of our method by obtaining state-of-the-art results on Ekman-6 and MovieNet. Another key contribution is a large-scale, emotion-labeled MIDI dataset for affective music generation. We then present an emotion-based MIDI generator, the first to condition on continuous emotional values rather than discrete categories, enabling nuanced music generation aligned with complex emotional content. To enhance temporal synchronization, we introduce a novel temporal boundary conditioning method, called "boundary offset encodings," aligning musical chords with scene changes. Combining video emotion classification, emotion-based music generation, and temporal boundary conditioning, EMSYNC emerges as a fully automatic video-based music generator. User studies show that it consistently outperforms existing methods in terms of music richness, emotional alignment, temporal synchronization, and overall preference, setting a new state-of-the-art in video-based music generation.

视频配乐情感生成连续情绪自动作曲

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。