用3D人脸追踪驱动扩散模型,实现高保真表情与身份一致的视频人脸动画。
ViDS: Video Diffusion Shader using 3D Face Tracking

- 基于参考图重建3DMM网格,结合驱动视频的姿势与表情参数动画化。
- 通过3DMM法向图提供密集几何线索,生成逼真且身份一致的肖像动画。
- 采用自回归扩散采样,延长生成时长并减少片段间不连续性,适合影视级人脸动画。
我们提出ViDS,一种利用3D人脸追踪进行表达丰富且身份保持的肖像视频生成方法。首先从参考图像重建特定身份的3DMM网格,并通过驱动视频中的表情与姿态参数进行动画化。利用3DMM法向图提供的密集几何线索,采用视频扩散模型作为神经着色器,合成逼真肖像动画并忠实保留参考图像的外观与身份。实验发现更精确的3DMM追踪可实现更细粒度的表情控制。我们还引入自回归扩散采样过程,扩展生成长度同时降低相邻片段间的不连续性。相较于依赖关键点条件或隐式运动潜变量的先前扩散方法,本方法在表达与姿态控制上更精细一致,且身份与外观保持更佳。详细的消融实验证明了设计选择的有效性。
原文摘要 · Abstract (English)
We introduce ViDS, a Video Diffusion Shader that leverages 3D face tracking for expressive and identity-preserving portrait animation. We first reconstruct the identity-specific 3DMM mesh from the reference image, and then animate it using expression and pose parameters from a driving video. Leveraging dense geometric cues from 3DMM normal maps, we employ a video diffusion model as a neural shader to synthesize lifelike portrait animations while preserving the appearance and identity of the reference image. We find that more accurate 3DMM tracking enables finer-grained expression control. We also introduce an autoregressive diffusion sampling process that extends generation beyond the model's native window while reducing discontinuities between adjacent clips. Compared with prior diffusion-based approaches for portrait animation that rely on landmark-based conditioning or implicit motion latents, our method achieves more detailed and consistent expression and pose control while faithfully preserving identity and appearance. Detailed ablation studies validate the effectiveness of our design choices. Project page: https://fusheng-ji.github.io/ViDS/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。