arXiv:2509.02141cs.GRcs.CV2025-09

首个实时高保真头像高斯可变形模型,细节更丰富,表情更自然。

GRMM: Real-Time High-Fidelity Gaussian Morphable Head Model with Learned Residuals

  • 用残差模块增强3DMM,提升皱纹、发际线等细节表现力
  • 75帧/秒实时渲染,支持全头覆盖和精细表情控制
  • 自建EXPRESS-50数据集,实现身份与表情精准解耦

3D可变形模型(3DMMs)可用于面部重建、动画与AR/VR中的可控几何与表情编辑,但传统基于PCA的网格模型在分辨率、细节和逼真度上受限。神经体素方法虽提升真实感,却难以满足交互需求。近期基于高斯泼溅(3DGS)的面部模型虽实现快速高质量渲染,但仍依赖网格3DMM先验控制表情,难以捕捉细微几何变化、表情特征及完整头部覆盖。本文提出GRMM,首个全头高斯3D可变形模型,通过添加残差几何与外观组件,在基础3DMM上恢复高频细节,如皱纹、皮肤纹理和发际线变化。该模型通过低维可解释参数(如身份形状、面部表情)实现解耦控制,并独立建模残差以捕捉超出基模型能力的个体与表情特异性细节。粗粒度解码器生成顶点级网格形变,细粒度解码器表示每个高斯的外观,轻量级CNN对光栅化图像进行后处理以增强真实感,整体保持75 FPS实时渲染。为学习一致且高保真的残差,我们构建了EXPRESS-50,首个包含50个身份、60种对齐表情的数据集,支持高斯3DMM中身份与表情的有效解耦。在单目3D人脸重建、新视角合成与表情迁移任务中,GRMM在保真度与表情准确性上超越现有最优方法,同时实现交互式实时性能。

原文摘要 · Abstract (English)

3D Morphable Models (3DMMs) enable controllable facial geometry and expression editing for reconstruction, animation, and AR/VR, but traditional PCA-based mesh models are limited in resolution, detail, and photorealism. Neural volumetric methods improve realism but remain too slow for interactive use. Recent Gaussian Splatting (3DGS) based facial models achieve fast, high-quality rendering but still depend solely on a mesh-based 3DMM prior for expression control, limiting their ability to capture fine-grained geometry, expressions, and full-head coverage. We introduce GRMM, the first full-head Gaussian 3D morphable model that augments a base 3DMM with residual geometry and appearance components, additive refinements that recover high-frequency details such as wrinkles, fine skin texture, and hairline variations. GRMM provides disentangled control through low-dimensional, interpretable parameters (e.g., identity shape, facial expressions) while separately modelling residuals that capture subject- and expression-specific detail beyond the base model's capacity. Coarse decoders produce vertex-level mesh deformations, fine decoders represent per-Gaussian appearance, and a lightweight CNN refines rasterised images for enhanced realism, all while maintaining 75 FPS real-time rendering. To learn consistent, high-fidelity residuals, we present EXPRESS-50, the first dataset with 60 aligned expressions across 50 identities, enabling robust disentanglement of identity and expression in Gaussian-based 3DMMs. Across monocular 3D face reconstruction, novel-view synthesis, and expression transfer, GRMM surpasses state-of-the-art methods in fidelity and expression accuracy while delivering interactive real-time performance.

3D人脸建模高斯泼溅实时渲染表情控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。