让视频扩散模型听懂音乐并跳舞,仅用单卡一天完成训练
MusicInfuser: Making Video Diffusion Listen and Dance
- 基于引导影响函数筛选可适配层,高效对齐音乐与视频生成
- 在有限数据下实现高质量舞蹈动作同步,支持长视频与新音乐
- 无需动作数据,适合创意生成与跨模态内容创作
我们提出 MusicInfuser,一种将预训练文本到视频扩散模型适配为音乐驱动舞蹈生成的方法。不同于从头训练多模态音频-视频或音频-运动模型,该方法通过一种基于引导启发的分层可适配性准则,选择可调整层,在仅使用少量特定数据集的情况下,显著降低训练成本并保留丰富先验知识。实验表明,MusicInfuser能有效弥合音乐与视频之间的鸿沟,生成响应动态、新颖且多样的舞蹈动作。框架在未见音乐、更长视频序列及非典型主体上均有良好泛化能力,且在一致性与同步性上优于基线模型。所有训练仅需单张GPU,一天内完成。
原文摘要 · Abstract (English)
We introduce MusicInfuser, an approach that aligns pre-trained text-to-video diffusion models to generate high-quality dance videos synchronized with specified music tracks. Rather than training a multimodal audio-video or audio-motion model from scratch, our method demonstrates how existing video diffusion models can be efficiently adapted to align with musical inputs. We propose a novel layer-wise adaptability criterion based on a guidance-inspired constructive influence function to select adaptable layers, significantly reducing training costs while preserving rich prior knowledge, even with limited, specialized datasets. Experiments show that MusicInfuser effectively bridges the gap between music and video, generating novel and diverse dance movements that respond dynamically to music. Furthermore, our framework generalizes well to unseen music tracks, longer video sequences, and unconventional subjects, outperforming baseline models in consistency and synchronization. All of this is achieved without requiring motion data, with training completed on a single GPU within a day.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。