arXiv:2601.05511cs.CV2026-01

用3D高斯点云实现可动画的高保真人脸替换,支持交互操控。

GaussianSwap: Animatable Video Face Swapping with 3D Gaussian Splatting

  • 基于3D高斯溅射构建可动态控制的人脸数字人
  • 在多个数据集上身份保留率超90%,时间一致性显著提升
  • 适合需要真实感人脸动画与交互的应用场景

我们提出GaussianSwap,一种新型视频人脸替换框架,通过目标视频构建基于3D高斯溅射的人脸数字人,并将源图像的身份信息迁移至该数字人。传统方法仅生成像素级面部表示,导致结果为无结构像素,无法进行动画或交互操作。本工作实现了从像素生成到高保真数字人构建的范式转变。框架首先预处理目标视频,提取FLAME参数、相机位姿和分割掩码;随后将3D高斯点云绑定至FLAME模型,实现跨帧动态面部控制。为确保身份保留,采用三种先进人脸识别模型联合构建复合身份嵌入,用于数字人微调。最后,在背景帧上渲染替换后的面部,生成最终视频。实验表明,GaussianSwap在身份保留(>90%)、视觉清晰度和时序一致性方面表现优异,且支持此前无法实现的交互式应用。

原文摘要 · Abstract (English)

We introduce GaussianSwap, a novel video face swapping framework that constructs a 3D Gaussian Splatting based face avatar from a target video while transferring identity from a source image to the avatar. Conventional video swapping frameworks are limited to generating facial representations in pixel-based formats. The resulting swapped faces exist merely as a set of unstructured pixels without any capacity for animation or interactive manipulation. Our work introduces a paradigm shift from conventional pixel-based video generation to the creation of high-fidelity avatar with swapped faces. The framework first preprocesses target video to extract FLAME parameters, camera poses and segmentation masks, and then rigs 3D Gaussian splats to the FLAME model across frames, enabling dynamic facial control. To ensure identity preserving, we propose an compound identity embedding constructed from three state-of-the-art face recognition models for avatar finetuning. Finally, we render the face-swapped avatar on the background frames to obtain the face-swapped video. Experimental results demonstrate that GaussianSwap achieves superior identity preservation, visual clarity and temporal consistency, while enabling previously unattainable interactive applications.

人脸替换3D高斯数字人动画生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。