arXiv:2506.18881cs.CVcs.MM2025-06被引 3

让视频自动跟着音乐节奏动起来,还能保持原画面内容。

Let Your Video Listen to Your Music!

  • 分两步:先对齐关键帧与音乐节拍,再用扩散模型补全中间帧。
  • 10分钟内完成特定视频适配,单张4090显卡即可实现。
  • 适合做音乐视频、宣传片的自动化剪辑,无需人工调速或拼接。

将视频中视觉运动的节奏与给定音乐轨道对齐,是多媒体制作中的实际需求,但在自主视频编辑领域仍属未充分探索的任务。有效的运动与音乐节拍对齐能显著提升观众参与度和视觉吸引力,尤其在音乐视频、宣传内容和电影剪辑中。现有方法通常依赖费力的手动剪辑、速度调整或启发式编辑技术。尽管部分生成模型可联合生成视频与音乐,但常将两者纠缠,限制了视频对音乐节拍的灵活对齐能力,同时难以保留完整视觉内容。本文提出一种新型高效框架MVAA(Music-Video Auto-Alignment),可自动编辑视频以匹配给定音乐节奏,同时保留原始视觉内容。为增强灵活性,我们将任务模块化为两步:首先将关键帧插入与音乐节拍对齐的时间点,随后使用帧条件扩散模型生成连贯的中间帧,保持原视频语义内容。为避免耗时的测试时训练,我们采用两阶段策略:先在小规模视频集上预训练修复模块以学习通用运动先验,再进行快速推理时微调以适应特定视频。该混合方法可在单张NVIDIA 4090 GPU上,仅用1个周期,在10分钟内完成适配。大量实验表明,本方法可实现高质量节拍对齐与视觉平滑性。

原文摘要 · Abstract (English)

Aligning the rhythm of visual motion in a video with a given music track is a practical need in multimedia production, yet remains an underexplored task in autonomous video editing. Effective alignment between motion and musical beats enhances viewer engagement and visual appeal, particularly in music videos, promotional content, and cinematic editing. Existing methods typically depend on labor-intensive manual cutting, speed adjustments, or heuristic-based editing techniques to achieve synchronization. While some generative models handle joint video and music generation, they often entangle the two modalities, limiting flexibility in aligning video to music beats while preserving the full visual content. In this paper, we propose a novel and efficient framework, termed MVAA (Music-Video Auto-Alignment), that automatically edits video to align with the rhythm of a given music track while preserving the original visual content. To enhance flexibility, we modularize the task into a two-step process in our MVAA: aligning motion keyframes with audio beats, followed by rhythm-aware video inpainting. Specifically, we first insert keyframes at timestamps aligned with musical beats, then use a frame-conditioned diffusion model to generate coherent intermediate frames, preserving the original video's semantic content. Since comprehensive test-time training can be time-consuming, we adopt a two-stage strategy: pretraining the inpainting module on a small video set to learn general motion priors, followed by rapid inference-time fine-tuning for video-specific adaptation. This hybrid approach enables adaptation within 10 minutes with one epoch on a single NVIDIA 4090 GPU using CogVideoX-5b-I2V as the backbone. Extensive experiments show that our approach can achieve high-quality beat alignment and visual smoothness.

视频生成节奏对齐扩散模型自动化剪辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。