将音乐生成舞蹈动作转化为多通道图像生成,提升节奏同步与动作连贯性。
Reframing Music-Driven 2D Dance Pose Generation as Multi-Channel Image Generation
- 把2D舞蹈动作编码为独热图像,用VAE压缩后由DiT模型生成。
- 在AIST++2D和大规模真实场景数据集上,动作与视频指标均优于现有方法。
- 适合需要高保真、长时序舞蹈生成的应用,如虚拟偶像与游戏动画。
近期的姿势到视频模型可将2D姿态序列转换为逼真的、身份保持的舞蹈视频,关键挑战在于从音乐生成时间连贯且节奏对齐的2D姿态,尤其在复杂、高方差的真实场景分布下。本文将音乐到舞蹈生成重构为音乐标记条件下的多通道图像合成问题:将2D姿态序列编码为独热图像,通过预训练图像VAE压缩,并采用DiT架构建模,从而继承现代文本到图像模型的结构与训练优势,更好地捕捉高方差2D姿态分布。在此框架基础上,提出(i)时间共享的时序索引机制,显式同步音乐标记与姿态潜变量的时间对应;(ii)参考姿态条件策略,保留个体身体比例与屏幕尺度,支持长时序分段拼接生成。在大型真实场景2D舞蹈语料库及校准后的AIST++2D基准测试中,各项姿态与视频空间指标及人类偏好评分均显著优于代表性方法,消融实验验证了表示、时序索引与参考条件的有效性。补充视频见 https://hot-dance.github.io
原文摘要 · Abstract (English)
Recent pose-to-video models can translate 2D pose sequences into photorealistic, identity-preserving dance videos, so the key challenge is to generate temporally coherent, rhythm-aligned 2D poses from music, especially under complex, high-variance in-the-wild distributions. We address this by reframing music-to-dance generation as a music-token-conditioned multi-channel image synthesis problem: 2D pose sequences are encoded as one-hot images, compressed by a pretrained image VAE, and modeled with a DiT-style backbone, allowing us to inherit architectural and training advances from modern text-to-image models and better capture high-variance 2D pose distributions. On top of this formulation, we introduce (i) a time-shared temporal indexing scheme that explicitly synchronizes music tokens and pose latents over time and (ii) a reference-pose conditioning strategy that preserves subject-specific body proportions and on-screen scale while enabling long-horizon segment-and-stitch generation. Experiments on a large in-the-wild 2D dance corpus and the calibrated AIST++2D benchmark show consistent improvements over representative music-to-dance methods in pose- and video-space metrics and human preference, and ablations validate the contributions of the representation, temporal indexing, and reference conditioning. See supplementary videos at https://hot-dance.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。