arXiv:2512.21905cs.CV2025-12被引 1

用扩散Transformer实现高保真长时序人像动画生成

High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer

  • 引入混合隐式引导与锐度因子,增强面部手部细节控制
  • 通过时序位置偏移融合模块,支持任意长度视频生成
  • 适合需要高精度动作与长时间动画的影视/虚拟人应用

扩散模型的进展显著推动了人像动画领域的发展。现有方法虽能生成短时或规则动作的时序一致结果,但在长时序视频生成及精细面部与手部细节合成方面仍存在挑战,限制了其在真实高质场景中的应用。为此,我们提出基于扩散Transformer(DiT)的框架,专注于生成高保真、长时序的人像动画视频。首先,设计了一组混合隐式引导信号和锐度引导因子,使框架能额外以面部与手部特征作为引导;其次,引入时序感知位置偏移融合模块,改进DiT主干的输入格式,提出位置偏移自适应模块,实现任意长度视频生成;最后,提出一种新型数据增强策略与骨骼对齐模型,降低不同身份间人体形态差异的影响。实验表明,本方法优于现有最先进方法,在高保真与长时序人像动画生成上均表现更优。

原文摘要 · Abstract (English)

Recent progress in diffusion models has significantly advanced the field of human image animation. While existing methods can generate temporally consistent results for short or regular motions, significant challenges remain, particularly in generating long-duration videos. Furthermore, the synthesis of fine-grained facial and hand details remains under-explored, limiting the applicability of current approaches in real-world, high-quality applications. To address these limitations, we propose a diffusion transformer (DiT)-based framework which focuses on generating high-fidelity and long-duration human animation videos. First, we design a set of hybrid implicit guidance signals and a sharpness guidance factor, enabling our framework to additionally incorporate detailed facial and hand features as guidance. Next, we incorporate the time-aware position shift fusion module, modify the input format within the DiT backbone, and refer to this mechanism as the Position Shift Adaptive Module, which enables video generation of arbitrary length. Finally, we introduce a novel data augmentation strategy and a skeleton alignment model to reduce the impact of human shape variations across different identities. Experimental results demonstrate that our method outperforms existing state-of-the-art approaches, achieving superior performance in both high-fidelity and long-duration human image animation.

人像动画扩散模型长视频生成细节重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。