仅用前后两张图秒级重建彩色3D人体,适合快速生成数字人。
Snap-Snap: Taking Two Images to Reconstruct 3D Human Gaussians in Milliseconds
- 基于基础重建模型设计新几何结构,实现稀疏视图下点云一致性预测
- 190毫秒完成1024x1024图像的完整人体重建,支持跨域数据集
- 可适配手机拍摄图像,降低3D数字人创建门槛
从稀疏视角重建3D人体是一项重要且有吸引力的任务。本文提出一个极具挑战性但实用的新任务:仅通过正面和背面两张图像重建人体,显著降低用户创建3D数字人的门槛。核心难点在于如何在输入视图重叠极少的情况下建立3D一致性并恢复缺失信息。为此,我们基于基础重建模型重新设计几何重建架构,在大量人体数据训练下仍能预测一致的点云。此外,引入增强算法补充缺失颜色信息,最终获得完整带色点云,并直接转换为3D高斯表示以提升渲染质量。实验表明,该方法在单张NVIDIA RTX 4090上仅需190毫秒即可完成1024x1024分辨率图像的完整人体重建,在THuman2.0及跨域数据集上表现达到当前最优水平。此外,方法对低质量手机拍摄图像也具备良好鲁棒性,大幅降低数据采集要求。演示与代码已公开于https://hustvl.github.io/Snap-Snap/。
原文摘要 · Abstract (English)
Reconstructing 3D human bodies from sparse views has been an appealing topic, which is crucial to broader the related applications. In this paper, we propose a quite challenging but valuable task to reconstruct the human body from only two images, i.e., the front and back view, which can largely lower the barrier for users to create their own 3D digital humans. The main challenges lie in the difficulty of building 3D consistency and recovering missing information from the highly sparse input. We redesign a geometry reconstruction model based on foundation reconstruction models to predict consistent point clouds even input images have scarce overlaps with extensive human data training. Furthermore, an enhancement algorithm is applied to supplement the missing color information, and then the complete human point clouds with colors can be obtained, which are directly transformed into 3D Gaussians for better rendering quality. Experiments show that our method can reconstruct the entire human in 190 ms on a single NVIDIA RTX 4090, with two images at a resolution of 1024x1024, demonstrating state-of-the-art performance on the THuman2.0 and cross-domain datasets. Additionally, our method can complete human reconstruction even with images captured by low-cost mobile devices, reducing the requirements for data collection. Demos and code are available at https://hustvl.github.io/Snap-Snap/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。