arXiv:2412.00397cs.CV2024-12ICCV被引 19

用2D姿态生成高质量人物动画,通过三维几何增强实现流畅动作

DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses

  • 从2D姿态推断3D几何信息,提升动作连贯性
  • 在TikTok-Dance5K数据集上达当前最佳效果
  • 适合视频生成与数字人动画开发者使用

本文提出DreamDance,一种仅需骨骼姿态序列即可驱动人物图像动画的新方法。现有方法或因缺乏3D信息导致质量不佳,或依赖复杂3D表示流程耗时。为解决此问题,DreamDance通过引入高效扩散模型,从2D姿态中增强3D几何线索,实现高质量动画生成。核心思想是捕捉从粗略骨架到精细几何、再到外观细节的多层级相关性,以增强引导信号,提升帧内一致性和帧间连贯性。我们构建了TikTok-Dance5K数据集,包含5000个高质量舞蹈视频,附带逐帧的人体姿态、深度图和法线图标注。进一步设计了互对齐几何扩散模型,用于生成细粒度深度图与法线图作为增强引导。最后,跨域控制器融合多层级引导,结合视频扩散模型实现有效动画生成。大量实验表明,该方法在人物图像动画任务中达到当前最优性能。

原文摘要 · Abstract (English)

In this work, we present DreamDance, a novel method for animating human images using only skeleton pose sequences as conditional inputs. Existing approaches struggle with generating coherent, high-quality content in an efficient and user-friendly manner. Concretely, baseline methods relying on only 2D pose guidance lack the cues of 3D information, leading to suboptimal results, while methods using 3D representation as guidance achieve higher quality but involve a cumbersome and time-intensive process. To address these limitations, DreamDance enriches 3D geometry cues from 2D poses by introducing an efficient diffusion model, enabling high-quality human image animation with various guidance. Our key insight is that human images naturally exhibit multiple levels of correlation, progressing from coarse skeleton poses to fine-grained geometry cues, and further from these geometry cues to explicit appearance details. Capturing such correlations could enrich the guidance signals, facilitating intra-frame coherency and inter-frame consistency. Specifically, we construct the TikTok-Dance5K dataset, comprising 5K high-quality dance videos with detailed frame annotations, including human pose, depth, and normal maps. Next, we introduce a Mutually Aligned Geometry Diffusion Model to generate fine-grained depth and normal maps for enriched guidance. Finally, a Cross-domain Controller incorporates multi-level guidance to animate human images effectively with a video diffusion model. Extensive experiments demonstrate that our method achieves state-of-the-art performance in animating human images.

人物动画扩散模型姿态驱动三维几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。