arXiv:2504.07594cs.MMcs.CV2025-04被引 6

提升视频动态建模,实现更精准的视频配乐生成

Extending Visual Dynamics for Video-to-Music Generation

  • 用简化光流编码器提取帧级动态特征,结合自注意力聚合
  • 通过动态特征扩展音乐令牌,解决音视频时间错位问题
  • 适合需要高质量音视频同步的影视、短视频创作场景

音乐能显著提升视频质量、互动性与情感共鸣,推动了视频到音乐生成的研究。尽管已有进展,现有方法在特定场景受限或低估视觉动态信息。为解决这些问题,本文提出 DyViM 框架,强化视频动态建模并缓解音视频表征的时间错位。具体而言,采用继承自光流方法的简化运动编码器提取逐帧动态特征,并通过自注意力模块在帧内聚合;这些动态特征被用于扩展现有音乐令牌以实现时序对齐。同时,高层语义通过交叉注意力机制传递,采用退火调优策略高效微调预训练音乐解码器,实现无缝适配。大量实验表明,DyViM 在性能上优于当前最先进(SOTA)方法。

原文摘要 · Abstract (English)

Music profoundly enhances video production by improving quality, engagement, and emotional resonance, sparking growing interest in video-to-music generation. Despite recent advances, existing approaches remain limited in specific scenarios or undervalue the visual dynamics. To address these limitations, we focus on tackling the complexity of dynamics and resolving temporal misalignment between video and music representations. To this end, we propose DyViM, a novel framework to enhance dynamics modeling for video-to-music generation. Specifically, we extract frame-wise dynamics features via a simplified motion encoder inherited from optical flow methods, followed by a self-attention module for aggregation within frames. These dynamic features are then incorporated to extend existing music tokens for temporal alignment. Additionally, high-level semantics are conveyed through a cross-attention mechanism, and an annealing tuning strategy benefits to fine-tune well-trained music decoders efficiently, therefore facilitating seamless adaptation. Extensive experiments demonstrate DyViM's superiority over state-of-the-art (SOTA) methods.

视频配乐动态建模跨模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。