arXiv:2502.10841cs.CV2025-02被引 40

用视频扩散Transformer实现高保真人脸动画,解决表情失真与身份错乱问题。

SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers

  • 基于视频扩散Transformer,引入表情感知条件模块提升动作迁移精度。
  • 在头像动画中保持身份一致且动作自然,生成视频视觉连贯性高。
  • 适合虚拟人、远程通信和数字内容生成场景,支持多样化身体比例。

我们提出SkyReels-A1,一个基于视频扩散Transformer的简洁高效框架,用于实现肖像图像的生动动画。现有方法在仅头部动画场景下仍存在身份失真、背景不稳和面部动作不自然等问题,且扩展至不同体型时易导致视觉不一致或动作异常。SkyReels-A1利用视频DiT强大的生成能力,提升面部运动迁移精度、身份保留度与时间连贯性。系统引入表达感知条件模块,实现由表情引导的关键点输入驱动的无缝视频合成;融合面部图文对齐模块,强化面部特征与运动轨迹的结合,增强身份一致性。此外,采用多阶段训练策略,逐步优化表情与动作间的关联,同时确保身份稳定重现。大量实证评估表明,该模型能生成视觉连贯且构图多样的结果,适用于虚拟形象、远程通信和数字媒体生成等场景。

原文摘要 · Abstract (English)

We present SkyReels-A1, a simple yet effective framework built upon video diffusion Transformer to facilitate portrait image animation. Existing methodologies still encounter issues, including identity distortion, background instability, and unrealistic facial dynamics, particularly in head-only animation scenarios. Besides, extending to accommodate diverse body proportions usually leads to visual inconsistencies or unnatural articulations. To address these challenges, SkyReels-A1 capitalizes on the strong generative capabilities of video DiT, enhancing facial motion transfer precision, identity retention, and temporal coherence. The system incorporates an expression-aware conditioning module that enables seamless video synthesis driven by expression-guided landmark inputs. Integrating the facial image-text alignment module strengthens the fusion of facial attributes with motion trajectories, reinforcing identity preservation. Additionally, SkyReels-A1 incorporates a multi-stage training paradigm to incrementally refine the correlation between expressions and motion while ensuring stable identity reproduction. Extensive empirical evaluations highlight the model's ability to produce visually coherent and compositionally diverse results, making it highly applicable to domains such as virtual avatars, remote communication, and digital media generation.

人脸动画视频生成扩散模型身份保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。