arXiv:2512.16893cs.CV2025-12

用单图快速生成高保真3D人脸动画,每秒超100帧。

Instant Expressive Gaussian Head Avatars at Over 100 FPS

  • 轻量局部融合机制,高效结合3D结构与表情信息。
  • 动画速度达107.31 FPS,比现有方法快3-4个数量级。
  • 无需参数化模型,适合实时数字人、远程呈现场景。

肖像动画近年来因视频扩散模型的进步而大幅提升质量,但这些2D方法常牺牲3D一致性与速度,限制了在数字孪生或远程通信等场景的应用。相比之下,基于3D表示(如神经辐射场或高斯溅射)的前馈式面部动画方法虽保证3D一致性并实现更快推理,却在表情细节上表现较差。本文解决这一三难困境(速度、3D一致性、表现力),提出一种前馈编码器管道,可将野外单图瞬间转换为3D一致、快速且富有表现力的可动画化表示。不同于以往计算量大的全局融合机制(如多层注意力),本设计采用高效轻量的局部融合策略,实现高表达力。此外,动画表示与3D人脸解耦,通过数据隐式学习运动,摆脱对预定义参数模型的依赖,从而突破动画能力限制。方法动画与姿态控制速度达107.31 FPS,相比最先进方法提速3-4个数量级,同时保持相当的动画质量,优于需以速度换质量或反之的设计。

原文摘要 · Abstract (English)

Portrait animation has witnessed tremendous quality improvements thanks to recent advances in video diffusion models. However, these 2D methods often compromise 3D consistency and speed, limiting their applicability in real-world scenarios, such as digital twins or telepresence. In contrast, 3D-aware feedforward facial animation methods -- built upon 3D representations, such as neural radiance fields or Gaussian splatting -- ensure 3D consistency and achieve faster inference speed, but come with inferior expression details. In this paper, we address this portrait animation trilemma (speed, 3D consistency, and expressiveness) and propose a pipeline that instantly converts an in-the-wild single image into a 3D-consistent, fast yet expressive animatable representation via a feed-forward encoder. Unlike previous computationally intensive global fusion mechanisms (e.g., multiple attention layers) for fusing 3D structural and animation information, our design employs an efficient lightweight local fusion strategy to achieve high animation expressivity. Furthermore, our animation representation is decoupled from the face's 3D representation and learns motion implicitly from data, eliminating the dependency on pre-defined parametric models that often constrain animation capabilities. Our method runs at 107.31 FPS for animation and pose control, representing a 3-4 order of magnitude speedup versus the state of the art while achieving comparable animation quality, thus surpassing alternative designs that trade speed for quality or vice versa.

3D人脸实时动画高斯溅射数字人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。