arXiv:2505.05589cs.CVcs.AI2025-05被引 1

用分层结构生成高精度且连贯的长时舞蹈动作,适合人机交互与虚拟演出。

ReactDance: Hierarchical Representation for High-Fidelity and Coherent Long-Form Reactive Dance Generation

  • 分层量化表示分离身体姿态与肢体细节,实现精准控制。
  • 块状并行生成使长序列输出效率提升3倍以上,同时保持时间连贯性。
  • 适合舞蹈生成、虚拟偶像、人机协作等场景,对动作细节要求高的应用。

反应式舞蹈生成(RDG)旨在根据领舞动作生成协同舞蹈,对增强人机交互和沉浸式数字娱乐具有重要意义。尽管在双人同步和动作-音乐对齐方面取得进展,仍面临精细空间互动与长期时间连贯性两大挑战。本文提出ReactDance,一种基于新型分层潜在空间的扩散框架,以解决这些时空难题。首先,为实现高保真空间表达与细粒度控制,提出分层有限标量量化(HFSQ),将粗略身体姿态与细微肢体动态解耦,通过分层引导机制分别精细调控。其次,为高效生成长序列并保证时间连贯性,提出块状局部上下文(BLC)非自回归采样策略:将序列分块并行合成,结合周期因果掩码与位置编码;通过密集滑动窗口训练强化局部时序上下文。大量实验表明,ReactDance在动作质量、长期连贯性及采样效率上均显著优于现有最优方法。项目主页:https://ripemangobox.github.io/ReactDance。

原文摘要 · Abstract (English)

Reactive dance generation (RDG), the task of generating a dance conditioned on a lead dancer's motion, holds significant promise for enhancing human-robot interaction and immersive digital entertainment. Despite progress in duet synchronization and motion-music alignment, two key challenges remain: generating fine-grained spatial interactions and ensuring long-term temporal coherence. In this work, we introduce \textbf{ReactDance}, a diffusion framework that operates on a novel hierarchical latent space to address these spatiotemporal challenges in RDG. First, for high-fidelity spatial expression and fine-grained control, we propose Hierarchical Finite Scalar Quantization (\textbf{HFSQ}). This multi-scale motion representation effectively disentangles coarse body posture from subtle limb dynamics, enabling independent and detailed control over both aspects through a layered guidance mechanism. Second, to efficiently generate long sequences with high temporal coherence, we propose Blockwise Local Context (\textbf{BLC}), a non-autoregressive sampling strategy. Departing from slow, frame-by-frame generation, BLC partitions the sequence into blocks and synthesizes them in parallel via periodic causal masking and positional encodings. Coherence across these blocks is ensured by a dense sliding-window training approach that enriches the representation with local temporal context. Extensive experiments show that ReactDance substantially outperforms state-of-the-art methods in motion quality, long-term coherence, and sampling efficiency. Project page: https://ripemangobox.github.io/ReactDance.

舞蹈生成扩散模型时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。