首个音乐驱动的多阶段视频生成框架,支持镜头运动与动作同步。
YingVideo-MV: Music-Driven Multi-Stage Video Generation
- 分阶段设计,融合音频语义分析与可解释镜头规划。
- 生成长视频时保持动作-音乐-镜头高度同步,提升连贯性。
- 适合音乐视频创作、虚拟演出等需要精确同步的应用场景。
尽管基于扩散模型的音频驱动形象视频生成已在长序列合成、自然音画同步和身份一致性方面取得显著进展,但音乐表演视频中镜头运动的生成仍缺乏探索。我们提出YingVideo-MV,首个用于音乐驱动长视频生成的级联框架。该方法结合音频语义分析、可解释的镜头规划模块(MV-Director)、时间感知扩散Transformer架构以及长序列一致性建模,实现从音频信号自动合成高质量音乐表演视频。我们通过收集网络数据构建了大规模《Music-in-the-Wild Dataset》以支持多样化、高质量的结果。针对现有长视频生成方法缺乏显式镜头运动控制的问题,引入相机适配模块,将相机姿态嵌入潜在噪声中。为增强长序列推理中片段间的连续性,进一步提出时间感知动态窗口策略,根据音频嵌入自适应调整去噪范围。综合基准测试表明,YingVideo-MV在生成连贯且富有表现力的音乐视频方面表现优异,实现了精确的音乐-动作-镜头同步。更多视频展示见项目页:https://giantailab.github.io/YingVideo-MV/
原文摘要 · Abstract (English)
While diffusion model for audio-driven avatar video generation have achieved notable process in synthesizing long sequences with natural audio-visual synchronization and identity consistency, the generation of music-performance videos with camera motions remains largely unexplored. We present YingVideo-MV, the first cascaded framework for music-driven long-video generation. Our approach integrates audio semantic analysis, an interpretable shot planning module (MV-Director), temporal-aware diffusion Transformer architectures, and long-sequence consistency modeling to enable automatic synthesis of high-quality music performance videos from audio signals. We construct a large-scale Music-in-the-Wild Dataset by collecting web data to support the achievement of diverse, high-quality results. Observing that existing long-video generation methods lack explicit camera motion control, we introduce a camera adapter module that embeds camera poses into latent noise. To enhance continulity between clips during long-sequence inference, we further propose a time-aware dynamic window range strategy that adaptively adjust denoising ranges based on audio embedding. Comprehensive benchmark tests demonstrate that YingVideo-MV achieves outstanding performance in generating coherent and expressive music videos, and enables precise music-motion-camera synchronization. More videos are available in our project page: https://giantailab.github.io/YingVideo-MV/ .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。