arXiv:2505.22977cs.CV2025-05被引 6

构建复杂人体动作动画数据集与评测基准,提升动态姿态生成质量

HyperMotionX: The Dataset and Benchmark with DiT-Based Pose-Guided Human Image Animation of Complex Motions

  • 基于DiT架构设计新模型,引入可学习频率缩放的RoPE模块增强低频空间特征
  • 在复杂动态动作下实现更高结构稳定性和外观一致性,显著优于现有方法
  • 开源高质量数据集与评测平台,推动复杂人体动作动画研究发展

扩散模型的进展显著提升了条件视频生成能力,尤其在姿态引导的人体图像动画任务中。尽管现有方法能在常规动作和静态场景中生成高保真、时间一致的动画序列,但在面对高度动态、非标准的人体动作时仍存在明显局限,且缺乏针对复杂人体动作动画的高质量评测基准。为此,我们提出一个简洁而强大的基于DiT的人体动画生成基线,并设计了空间低频增强型RoPE模块,通过引入可学习的频率缩放,选择性增强低频空间特征建模。此外,我们推出了Open-HyperMotionX数据集与HyperMotionX评测基准,提供高质量人体姿态标注和精选视频片段,用于评估和改进姿态引导的人体图像动画模型在复杂动作条件下的表现。大量实验表明,所提方法在高度动态人体动作序列中显著提升了结构稳定性和外观一致性。代码、模型权重及数据集已公开于https://vivocameraresearch.github.io/hypermotion/

原文摘要 · Abstract (English)

Recent advances in diffusion models have significantly improved conditional video generation, particularly in the pose-guided human image animation task. Although existing methods are capable of generating high-fidelity and time-consistent animation sequences in regular motions and static scenes. However there are still obvious limitations when facing complex human body motions that contain highly dynamic, non-standard motions, and the lack of a high-quality benchmark for evaluation of complex human motion animations. To address this challenge, we propose a concise yet powerful DiT-based human animation generation baseline and design spatial low-frequency enhanced RoPE, a novel module that selectively enhances low-frequency spatial feature modeling by introducing learnable frequency scaling. Furthermore, we introduce the Open-HyperMotionX Dataset and HyperMotionX Bench, which provide high-quality human pose annotations and curated video clips for evaluating and improving pose-guided human image animation models under complex human motion conditions. Our method significantly improves structural stability and appearance consistency in highly dynamic human motion sequences. Extensive experiments demonstrate the effectiveness of our dataset and proposed approach in advancing the generation quality of complex human motion image animations. The codes, model weights, and dataset have been made publicly available at https://vivocameraresearch.github.io/hypermotion/

人体动画扩散模型姿态引导数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。