让音乐驱动舞蹈生成,保持动作自然且人物形象不变。
FlowDance: Music-Driven Dance Video Generation with Parallel Pose and RGB Streams

- 并行处理姿态与图像流,分别建模动作和外观。
- 在去噪过程中动态调整姿态引导,长期保持身份特征一致。
- 适合需要高质量舞蹈视频生成的研究与创作人员。
音乐驱动的舞蹈视频合成旨在根据给定音乐片段动画化参考人物。该任务极具挑战性,因需同时学习音乐到动作的对应关系、保持身份的人体动画、时间连贯性以及视觉真实的视频生成。本文提出 FlowDance 框架,通过并行的姿态流与 RGB 流,将显式动作建模与参考保留的视觉合成相结合。引入时间步感知的姿态注入机制,以在去噪过程中自适应地提供结构引导;采用持续身份注入策略,在长视频中保持参考外观的一致性。为支持该任务,构建了一个精选的高分辨率真实场景舞蹈视频数据集,包含同步音乐、RGB 视频、3D 人体运动、相机参数及投影 2D 姿态标注。大量实验表明,FlowDance 在舞蹈动作生成与音乐驱动视频合成方面均表现优异。
原文摘要 · Abstract (English)
Music-driven dance video synthesis aims to animate a reference person according to a given music clip. The task is challenging because it requires a model to jointly learn music-to-motion correspondence, identity-preserving human animation, temporal coherence, and visually realistic video generation. We present FlowDance, a music-driven dance video generation framework that integrates explicit motion modeling with reference-preserving visual synthesis through parallel pose and RGB streams. We further introduce timestep-aware pose injection to adapt structural guidance across denoising steps and persistent identity injection to preserve the reference appearance over long video. To support this task, we further build a popularity-curated, high-resolution in-the-wild dance video dataset with synchronized music, RGB videos, 3D body motion, camera parameters, and projected 2D pose annotations. Extensive experiments show that FlowDance achieves strong performance in both dance motion generation and music-driven dance video synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。