用多视角图像高效重建高精度3D人脸,支持大规模扩展。
Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures

- 通过编码器-解码器压缩多视角图像为紧凑隐空间表示
- 在10,000+人规模数据集上实现顶尖重建质量
- 可生成新身份或驱动表情,无需测试时优化
我们提出 HeadsUp,一种可扩展的前馈方法,用于从大规模多相机设置中重建高质量3D高斯人脸。该方法采用高效的编码器-解码器架构,将输入视角压缩为紧凑的潜在表示,并解码为锚定于中性头模板的UV参数化3D高斯分布。该UV表示使3D高斯数量与输入图像数量和分辨率解耦,从而支持使用大量高分辨率图像进行训练。我们在一个包含超过10,000名受试者的内部数据集上训练和评估模型,该数据集规模比现有多视角人脸数据集大一个数量级。HeadsUp 实现了最先进的重建质量,并能泛化到未见身份而无需测试时优化。我们对模型在身份、视角和模型容量上的缩放行为进行了广泛分析,揭示了质量-计算权衡的实际见解。最后,我们展示了潜在空间的强大能力,包括生成新3D身份和使用表情混合形状动画3D头部。
原文摘要 · Abstract (English)
We propose HeadsUp, a scalable feed-forward method for reconstructing high-quality 3D Gaussian heads from large-scale multi-camera setups. Our method employs an efficient encoder-decoder architecture that compresses input views into a compact latent representation. This latent representation is then decoded into a set of UV-parameterized 3D Gaussians anchored to a neutral head template. This UV representation decouples the number of 3D Gaussians from the number and resolution of input images, enabling training with many high-resolution input views. We train and evaluate our model on an internal dataset with more than 10,000 subjects, which is an order of magnitude larger than existing multi-view human head datasets. HeadsUp achieves state-of-the-art reconstruction quality and generalizes to novel identities without test-time optimization. We extensively analyze the scaling behavior of our model across identities, views, and model capacity, revealing practical insights for quality-compute trade-offs. Finally, we highlight the strength of our latent space by showcasing two downstream applications: generating novel 3D identities and animating the 3D heads with expression blendshapes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。