arXiv:2507.23143cs.CV2025-07ICLR被引 33

用1维隐向量控制表情,实现零样本人脸动画生成

X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention

  • 通过交叉注意力机制将驱动视频动作编码为1维运动隐向量
  • 在多个数据集上实现比现有方法更逼真的表情表现和身份保留
  • 适合需要高质量人脸重演的视频生成与数字人研究者

我们提出X-NeMo,一种基于扩散模型的零样本肖像动画新方法,可利用不同个体的驱动视频中的面部动作,动画化静态肖像。针对以往方法存在的身份泄露和难以捕捉细微及极端表情的问题,我们设计了端到端训练框架,从驱动图像中提取无身份依赖的一维运动隐表示,并在图像生成中通过交叉注意力控制运动。该隐式运动表示在多样化视频数据集上端到端学习,无需预训练运动检测器。通过双判别器架构监督运动潜变量学习,并结合空间与色彩增强,进一步提升表达力并解耦运动与身份特征。通过将驱动运动嵌入一维潜在向量,而非使用加性空间引导,我们的设计避免了结构线索从驱动条件传递至扩散主干,显著缓解身份泄露问题。大量实验表明,X-NeMo优于当前最优基线,在保持身份相似性的同时生成高度逼真的动画效果。代码与模型已开源。

原文摘要 · Abstract (English)

We propose X-NeMo, a novel zero-shot diffusion-based portrait animation pipeline that animates a static portrait using facial movements from a driving video of a different individual. Our work first identifies the root causes of the key issues in prior approaches, such as identity leakage and difficulty in capturing subtle and extreme expressions. To address these challenges, we introduce a fully end-to-end training framework that distills a 1D identity-agnostic latent motion descriptor from driving image, effectively controlling motion through cross-attention during image generation. Our implicit motion descriptor captures expressive facial motion in fine detail, learned end-to-end from a diverse video dataset without reliance on pretrained motion detectors. We further enhance expressiveness and disentangle motion latents from identity cues by supervising their learning with a dual GAN decoder, alongside spatial and color augmentations. By embedding the driving motion into a 1D latent vector and controlling motion via cross-attention rather than additive spatial guidance, our design eliminates the transmission of spatial-aligned structural clues from the driving condition to the diffusion backbone, substantially mitigating identity leakage. Extensive experiments demonstrate that X-NeMo surpasses state-of-the-art baselines, producing highly expressive animations with superior identity resemblance. Our code and models are available for research.

人脸动画扩散模型表情生成零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。