arXiv:2508.09959cs.CV2025-08被引 5

LIA-X通过可解释的潜空间运动控制,实现人脸动画的精准编辑。

LIA-X: Interpretable Latent Portrait Animator

  • 基于稀疏运动词典的线性潜空间导航,解耦面部动态因素。
  • 在多个基准上超越现有方法,在自重演与跨重演任务中表现优异。
  • 支持用户精细操控表情与姿态,适合视频编辑与3D人脸动画应用。

我们提出LIA-X,一种新型可解释的人脸动画生成模型,能够将驱动视频中的面部动态精确迁移至源人脸图像,并支持细粒度控制。LIA-X是一种自编码器,将运动迁移建模为潜空间中运动码的线性导航。其核心创新在于引入稀疏运动词典,使模型能将面部动态解耦为可解释的因素。不同于以往的‘形变-渲染’范式,该可解释性支持高效的‘编辑-形变-渲染’策略,实现对源图像中细粒度面部语义的精确操控,有效缩小源图像与驱动视频在姿态和表情上的初始差异。此外,我们在大规模数据集上成功训练了一个约10亿参数的大型模型。实验表明,本方法在多个基准上的自重演与跨重演任务中均优于现有方法。LIA-X的可解释性与可控性也使其适用于精细用户引导的图像与视频编辑,以及3D感知的人脸视频操作。

原文摘要 · Abstract (English)

We introduce LIA-X, a novel interpretable portrait animator designed to transfer facial dynamics from a driving video to a source portrait with fine-grained control. LIA-X is an autoencoder that models motion transfer as a linear navigation of motion codes in latent space. Crucially, it incorporates a novel Sparse Motion Dictionary that enables the model to disentangle facial dynamics into interpretable factors. Deviating from previous 'warp-render' approaches, the interpretability of the Sparse Motion Dictionary allows LIA-X to support a highly controllable 'edit-warp-render' strategy, enabling precise manipulation of fine-grained facial semantics in the source portrait. This helps to narrow initial differences with the driving video in terms of pose and expression. Moreover, we demonstrate the scalability of LIA-X by successfully training a large-scale model with approximately 1 billion parameters on extensive datasets. Experimental results show that our proposed method outperforms previous approaches in both self-reenactment and cross-reenactment tasks across several benchmarks. Additionally, the interpretable and controllable nature of LIA-X supports practical applications such as fine-grained, user-guided image and video editing, as well as 3D-aware portrait video manipulation.

人脸动画可解释性潜空间控制视频编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。