arXiv:2604.13171cs.CV2026-04

仅用几张照片和摄像头视频,就能重建高保真3D人脸动画

3DRealHead: Few-Shot Detailed Head Avatar

论文配图:3DRealHead: Few-Shot Detailed Head Avatar
图 1 · 摘自论文原文
  • 用单视角视频提取表情信号,结合风格化U-Net生成3D高斯点云
  • 在NeRSemble数据集上训练先验,实现少样本下身份与表情精准还原
  • 新增嘴部特征捕捉,显著提升对复杂表情的表达能力

人脸是交流的核心。为实现沉浸式数字人呈现,需忠实还原个体特征与细腻表情。现有3D头像方法虽有多视角数据或学习先验,仍难以准确再现身份与表情,尤其在口腔、牙齿等高度个性化的区域。此外,多数方法依赖3DMM表达控制,限制了表现力。为此,我们提出3DRealHead,一种少样本头像重建方法。用户仅需拍摄少量照片,再通过消费级摄像头采集视频,即可重建并驱动3D头像。该方法基于在NeRSemble数据集上训练的风格化U-Net先验,将3D头像表示为可渲染的新视角高斯点云。动画时,网络同时接收3DMM表情信号与驱动视频中提取的嘴部特征,从而还原3DMM无法表示的复杂表情,显著提升真实感。

原文摘要 · Abstract (English)

The human face is central to communication. For immersive applications, the digital presence of a person should mirror the physical reality, capturing the users idiosyncrasies and detailed facial expressions. However, current 3D head avatar methods often struggle to faithfully reproduce the identity and facial expressions, despite having multi-view data or learned priors. Learning priors that capture the diversity of human appearances, especially, for regions with highly person-specific features, like the mouth and teeth region is challenging as the underlying training data is limited. In addition, many of the avatar methods are purely relying on 3D morphable model-based expression control which strongly limits expressivity. To address these challenges, we are introducing 3DRealHead, a few-shot head avatar reconstruction method with a novel expression control signal that is extracted from a monocular video stream of the subject. Specifically, the subject can take a few pictures of themselves, recover a 3D head avatar and drive it with a consumer-level webcam. The avatar reconstruction is enabled via a novel few-shot inversion process of a 3D human head prior which is represented as a Style U-Net that emits 3D Gaussian primitives which can be rendered under novel views. The prior is learned on the NeRSemble dataset. For animating the avatar, the U-Net is conditioned on 3DMM-based facial expression signals, as well as features of the mouth region extracted from the driving video. These additional mouth features allow us to recover facial expressions that cannot be represented by the 3DMM leading to a higher expressivity and closer resemblance to the physical reality.

3D人脸重建少样本学习表情驱动高斯点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。