arXiv:2512.11645cs.CV2025-12被引 1

通过解耦表情、姿态和视角,实现高精度可控人脸动画生成。

FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint

论文配图:FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint
图 1 · 摘自论文原文
  • 分离控制表情、头部动作与摄像机视角,提升动画自由度。
  • 在真实感、表现力与视角一致性上超越现有方法。
  • 适合需要精细控制的人像视频生成场景。

我们提出FactorPortrait,一种用于可控人脸动画的视频扩散方法,能够从分离的控制信号(面部表情、头部运动、相机视角)中生成逼真视频。给定一张人脸图像、一段驱动视频和相机轨迹,该方法可将驱动视频中的表情与头部动作迁移至目标人脸,同时支持任意视角的新视图合成。利用预训练图像编码器从驱动视频中提取面部表情隐向量作为动画控制信号,这些隐向量隐式捕捉细微的表情动态,且身份与姿态信息被解耦,通过提出的表情控制器高效注入视频扩散变换器。对于相机与头部姿态控制,采用从3D身体网格追踪中渲染的Plücker射线图和法线图。为训练模型,我们构建了一个大规模合成数据集,包含多样化的相机视角、头部姿态与表情动态组合。大量实验表明,该方法在真实感、表现力、控制精度与视角一致性方面均优于现有方法。

原文摘要 · Abstract (English)

We introduce FactorPortrait, a video diffusion method for controllable portrait animation that enables lifelike synthesis from disentangled control signals of facial expressions, head movement, and camera viewpoints. Given a single portrait image, a driving video, and camera trajectories, our method animates the portrait by transferring facial expressions and head movements from the driving video while simultaneously enabling novel view synthesis from arbitrary viewpoints. We utilize a pre-trained image encoder to extract facial expression latents from the driving video as control signals for animation generation. Such latents implicitly capture nuanced facial expression dynamics with identity and pose information disentangled, and they are efficiently injected into the video diffusion transformer through our proposed expression controller. For camera and head pose control, we employ Plücker ray maps and normal maps rendered from 3D body mesh tracking. To train our model, we curate a large-scale synthetic dataset containing diverse combinations of camera viewpoints, head poses, and facial expression dynamics. Extensive experiments demonstrate that our method outperforms existing approaches in realism, expressiveness, control accuracy, and view consistency.

人脸动画视频生成扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。