arXiv:2505.03351cs.CV2025-05ICCV被引 22

仅用一张图快速生成可动的高保真上半身3D虚拟人

GUAVA: Generalizable Upper Body 3D Gaussian Avatar

  • 基于单图逆纹理映射与投影采样重建上半身高斯点云
  • 0.1秒完成重建,渲染质量显著优于现有方法
  • 支持实时动画,适合虚拟主播、游戏等场景

从单张图像高保真、可动画化地重建包含丰富面部和手部动作的上半身3D虚拟人,具有广泛的应用潜力。传统方法通常需要多视角或单视角视频,并依赖个体身份训练,过程复杂耗时。且受限于SMPLX的表达能力,多数方法仅关注身体动作而难以捕捉精细面部表情。为此,我们首先提出一个增强型人体模型(EHM)以提升面部表现力,并开发了精准的追踪方法。在此基础上,提出GUAVA——首个实现快速可动画化上半身3D高斯虚拟人重建的框架。通过逆纹理映射与投影采样技术,从单张图像推断上半身高斯点;再经神经优化器精修渲染图像。实验表明,GUAVA在渲染质量上显著超越现有方法,重建时间低于0.1秒,支持实时动画与渲染。

原文摘要 · Abstract (English)

Reconstructing a high-quality, animatable 3D human avatar with expressive facial and hand motions from a single image has gained significant attention due to its broad application potential. 3D human avatar reconstruction typically requires multi-view or monocular videos and training on individual IDs, which is both complex and time-consuming. Furthermore, limited by SMPLX's expressiveness, these methods often focus on body motion but struggle with facial expressions. To address these challenges, we first introduce an expressive human model (EHM) to enhance facial expression capabilities and develop an accurate tracking method. Based on this template model, we propose GUAVA, the first framework for fast animatable upper-body 3D Gaussian avatar reconstruction. We leverage inverse texture mapping and projection sampling techniques to infer Ubody (upper-body) Gaussians from a single image. The rendered images are refined through a neural refiner. Experimental results demonstrate that GUAVA significantly outperforms previous methods in rendering quality and offers significant speed improvements, with reconstruction times in the sub-second range (0.1s), and supports real-time animation and rendering.

3D虚拟人高斯溅射单图重建实时动画

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。