arXiv:2502.04847cs.CV2025-02被引 46

HumanDiT用扩散变压器生成长序列人体动作视频,细节更真实、一致性更强。

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

  • 基于扩散变压器架构,支持多分辨率和可变长度序列生成。
  • 在1.4万小时高质量数据上训练,生成视频细节清晰、姿态准确。
  • 适合需要长时序人体动作生成的场景,如影视动画、虚拟人应用。

人体动作视频生成已取得显著进展,但现有方法在长序列和复杂动作中仍难以准确渲染手部、面部等细节,且依赖固定分辨率,难以保持视觉一致性。为此,我们提出HumanDiT,一种基于扩散变压器(DiT)的姿势引导框架,在包含14,000小时高质量视频的大规模野数据集上训练,以生成高保真、细粒度人体渲染的视频。具体而言:(i) HumanDiT基于DiT架构,支持多种视频分辨率和可变序列长度,有利于长序列视频生成;(ii) 引入前缀潜在参考策略,保持长序列中的人物个性化特征。推理时,HumanDiT利用Keypoint-DiT生成后续姿态序列,实现从静态图像或已有视频延续生成;同时通过姿态适配器支持给定姿态序列的转移。大量实验表明,其在多样场景下均能生成长时序、姿态精确的视频,表现优异。

原文摘要 · Abstract (English)

Human motion video generation has advanced significantly, while existing methods still struggle with accurately rendering detailed body parts like hands and faces, especially in long sequences and intricate motions. Current approaches also rely on fixed resolution and struggle to maintain visual consistency. To address these limitations, we propose HumanDiT, a pose-guided Diffusion Transformer (DiT)-based framework trained on a large and wild dataset containing 14,000 hours of high-quality video to produce high-fidelity videos with fine-grained body rendering. Specifically, (i) HumanDiT, built on DiT, supports numerous video resolutions and variable sequence lengths, facilitating learning for long-sequence video generation; (ii) we introduce a prefix-latent reference strategy to maintain personalized characteristics across extended sequences. Furthermore, during inference, HumanDiT leverages Keypoint-DiT to generate subsequent pose sequences, facilitating video continuation from static images or existing videos. It also utilizes a Pose Adapter to enable pose transfer with given sequences. Extensive experiments demonstrate its superior performance in generating long-form, pose-accurate videos across diverse scenarios.

动作生成扩散模型视频生成姿态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。