arXiv:2603.28003cs.CV2026-03AAAI

从单视频生成高保真个性化3D人脸,分离结构与细节特征。

DipGuava: Disentangling Personalized Gaussian Features for 3D Head Avatars from Monocular Video

论文配图:DipGuava: Disentangling Personalized Gaussian Features for 3D Head Avatars from Monocular Video
图 1 · 摘自论文原文
  • 分两阶段学习:先建全局结构外观,再补个性化细节。
  • 重建精度和视觉质量超越现有方法,尤其在皱纹等细节上表现优异。
  • 适合需要真实感人脸动画的虚拟人、影视制作场景。

尽管近期的3D人脸动画方法尝试捕捉面部动态,但往往无法保留个性化细节,限制了真实感与表现力。为此,我们提出DipGuava(解耦且个性化的高斯纹理头像),一种从单视角视频生成个性化3D高斯头像的新方法。DipGuava首次显式地将面部外观解耦为两个互补成分,采用结构化的两阶段训练流程,显著降低学习歧义并提升重建保真度。第一阶段学习由几何驱动的基础外观,捕捉整体面部结构及粗粒度的表情变化;第二阶段预测第一阶段未覆盖的个性化残差细节,包括高频成分及非线性变化特征,如皱纹和细微皮肤形变。通过动态外观融合机制,在形变后整合残差细节,确保空间与语义对齐。该解耦设计使DipGuava能生成逼真且保持身份一致的头像,在大量实验中持续优于现有方法,无论在视觉质量还是定量指标上均表现更优。

原文摘要 · Abstract (English)

While recent 3D head avatar creation methods attempt to animate facial dynamics, they often fail to capture personalized details, limiting realism and expressiveness. To fill this gap, we present DipGuava (Disentangled and Personalized Gaussian UV Avatar), a novel 3D Gaussian head avatar creation method that successfully generates avatars with personalized attributes from monocular video. DipGuava is the first method to explicitly disentangle facial appearance into two complementary components, trained in a structured two-stage pipeline that significantly reduces learning ambiguity and enhances reconstruction fidelity. In the first stage, we learn a stable geometry-driven base appearance that captures global facial structure and coarse expression-dependent variations. In the second stage, the personalized residual details not captured in the first stage are predicted, including high-frequency components and nonlinearly varying features such as wrinkles and subtle skin deformations. These components are fused via dynamic appearance fusion that integrates residual details after deformation, ensuring spatial and semantic alignment. This disentangled design enables DipGuava to generate photorealistic, identity-preserving avatars, consistently outperforming prior methods in both visual quality and quantitativeperformance, as demonstrated in extensive experiments.

3D人脸高斯渲染个性化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。