arXiv:2506.06645cs.CV2025-06被引 6

用可学习人体先验,让单目视频快速生成逼真虚拟人像。

Parametric Gaussian Human Model: Generalizable Prior for Efficient and Realistic Human Avatar Modeling

  • 将人体几何与外观编码为可学习的隐空间特征图。
  • 20分钟内完成单目视频重建,质量媲美长时间优化方法。
  • 适合需快速生成真人数字形象的VR/远程参会场景。

逼真的可动画人体虚拟形象是虚拟/增强现实、远程通信和数字娱乐的关键。尽管3D高斯点云(3DGS)技术显著提升了渲染质量和效率,现有方法仍面临两大挑战:需耗时进行个体优化,且在稀疏单目输入下泛化能力差。本文提出参数化高斯人体模型(PGHM),通过将人体先验融入3DGS,实现从单目视频快速、高保真地重建人体虚拟形象。PGHM包含两个核心组件:(1) 采用UV对齐的潜在身份图,将个体特异性几何与外观紧凑编码为可学习特征张量;(2) 解耦的多头U-Net,通过条件解码器分离静态、姿态依赖和视角依赖成分,预测高斯属性。该设计在复杂姿态和视角下仍保持鲁棒渲染质量,且无需多视角数据或长时间优化即可实现高效主体适配。实验表明,相比从零优化的方法,PGHM每主体仅需约20分钟即可生成视觉质量相当的虚拟形象,展现出其在真实单目虚拟人创建中的实用性。

原文摘要 · Abstract (English)

Photorealistic and animatable human avatars are a key enabler for virtual/augmented reality, telepresence, and digital entertainment. While recent advances in 3D Gaussian Splatting (3DGS) have greatly improved rendering quality and efficiency, existing methods still face fundamental challenges, including time-consuming per-subject optimization and poor generalization under sparse monocular inputs. In this work, we present the Parametric Gaussian Human Model (PGHM), a generalizable and efficient framework that integrates human priors into 3DGS for fast and high-fidelity avatar reconstruction from monocular videos. PGHM introduces two core components: (1) a UV-aligned latent identity map that compactly encodes subject-specific geometry and appearance into a learnable feature tensor; and (2) a disentangled Multi-Head U-Net that predicts Gaussian attributes by decomposing static, pose-dependent, and view-dependent components via conditioned decoders. This design enables robust rendering quality under challenging poses and viewpoints, while allowing efficient subject adaptation without requiring multi-view capture or long optimization time. Experiments show that PGHM is significantly more efficient than optimization-from-scratch methods, requiring only approximately 20 minutes per subject to produce avatars with comparable visual quality, thereby demonstrating its practical applicability for real-world monocular avatar creation.

人体建模3D高斯单目重建虚拟形象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。