arXiv:2607.09581cs.CVcs.SD2026-07被引 1

突破分钟级舞蹈视频生成瓶颈,实现节奏同步与动作连贯

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

论文配图:Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation
图 1 · 摘自论文原文
  • 分层架构先规划关键帧再细化时序,提升长程一致性
  • 生成超1分钟720p/30fps视频,避免动作重复与漂移
  • 支持五类舞种,适配音频+文本双条件输入

直接从音乐生成长时间、高画质且节奏同步的舞蹈视频仍面临重大挑战,主要源于当前扩散模型在20秒以上时的时序限制。现有方法或依赖中间3D骨骼,或采用端到端视频合成,扩展至长时域时易出现时间漂移、身份不一致和重复动作。为此,我们提出一种新型分层框架,实现分钟级连贯的音乐到舞蹈生成。该方法将过程解耦为全局关键帧规划与局部时序精炼,利用全曲音乐上下文保障长程一致性。核心创新包括:通过时序映射的RoPE嵌入实现动态帧率自适应以精准对齐;基于光流的损失函数增强运动连续性;运动速度控制机制在快速动作中保持高保真细节。大量实验表明,该框架突破传统时长限制,可生成稳定、超过一分钟、720p/30fps的视频,具备优异的时间稳定性。同时,模型在五种不同舞蹈风格下表现稳健,可接受音频与文本双重提示,建立长时舞蹈视频合成新基准。

原文摘要 · Abstract (English)

Generating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typically fail beyond 20 seconds. Existing approaches, whether they rely on intermediate 3D skeletons or on end-to-end video synthesis, suffer from temporal drift, identity inconsistency, and repetitive motion patterns when extended to longer horizons. To address these limitations, we propose a novel hierarchical framework for minute-scale coherent music-to-dance generation. Our method decouples the process into global keyframe planning and local temporal refinement, leveraging full-track musical context to ensure long-range coherence. Key innovations include dynamic frame rate adaptation via time-mapped RoPE embeddings for precise alignment, an optical-flow-based loss function to enhance motion continuity, and motion-speed control to preserve high-fidelity details during rapid movements. Extensive experiments demonstrate that our framework surpasses the conventional duration barrier, generating stable, 720p/30fps videos exceeding one minute with superior temporal stability. Furthermore, the model exhibits robust versatility across five distinct dance genres, conditioned on both audio and textual prompts, establishing a new state-of-the-art in coherent, long-form dance video synthesis.

舞蹈生成扩散模型长时序多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。