用少量摄像头图像生成高保真动态虚拟人,支持任意新人物泛化。
GIGA: Generalizable Sparse Image-driven Gaussian Humans
- 基于多视角图像构建3D高斯点云,通过多头网络预测纹理。
- 训练可扩展至数千人,保持高保真与动态外观生成能力。
- 适合虚拟现实、数字人制作等需快速建模的场景。
从少量RGB摄像头驱动高质量、逼真的全身虚拟人类是虚拟现实技术中的关键挑战。理想的解决方案应能泛化到任意人物,仅需单视角或稀疏多视角视频即可生成自由视角的逼真渲染。现有方法难以扩展至大规模数据集,导致多样性不足且保真度有限。为此,我们提出GIGA,一种新型通用全身体三维虚拟人模型,可由1至4个输入视角及追踪的身体模板生成自由视角的高保真渲染。核心创新为多头UNet架构,将单视图或多视图累积的近似RGB纹理作为输入,预测位于人体网格上的3D高斯原语(以2D texels表示)。训练阶段可扩展至数千名个体,同时保持高保真度与动态外观合成能力。在身份泛化与视觉保真度上显著优于现有方法。
原文摘要 · Abstract (English)
Driving a high-quality and photorealistic full-body virtual human from a few RGB cameras is a challenging problem that has become increasingly relevant with emerging virtual reality technologies. A promising solution to democratize such technology would be a generalizable method that takes sparse multi-view images of any person and then generates photoreal free-view renderings of them. However, the state-of-the-art approaches are not scalable to very large datasets and, thus, lack diversity and photorealism. To address this problem, we propose GIGA, a novel, generalizable full-body model for rendering photoreal humans in free viewpoint, driven by a single-view or sparse multi-view video. Notably, GIGA can scale training to a few thousand subjects while maintaining high photorealism and synthesizing dynamic appearance. At the core, we introduce a MultiHeadUNet architecture, which takes an approximate RGB texture accumulated from a single or multiple sparse views and predicts 3D Gaussian primitives represented as 2D texels on top of a human body mesh. At test time, our method performs novel view synthesis of a virtual 3D Gaussian-based human from 1 to 4 input views and a tracked body template for unseen identities. Our method excels over prior works by a significant margin in terms of identity generalization capability and photorealism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。