用单目视频快速重建高精度3D人像,训练速度比现有方法快110倍。
Efficient Neural Implicit Representation for 3D Human Reconstruction
- 融合HuMoR、Instant-NGP与Fast-SNARF,实现高效姿态与外观重建
- 30秒训练即可生成可用视觉结果,全程仅需数分钟
- 适合实时交互、元宇宙等对效率要求高的应用场景
高保真数字人像在互动远程呈现、AR/VR、3D图形及不断发展的元宇宙中需求日益增长。尽管传统方法在小空间内表现良好,但通常依赖昂贵硬件且计算成本高。本文提出HumanAvatar,一种从单目视频高效重建精确人像的新方法。核心在于融合预训练的HuMoR(擅长人体运动估计)、前沿神经辐射场技术Instant-NGP以及先进刚性模型Fast-SNARF,显著提升重建质量与速度。通过引入姿态敏感的空间压缩技术,系统在渲染质量与计算效率间取得最优平衡。在合成与真实单目视频上的详尽实验表明,HumanAvatar在质量上持续优于或等于当前领先方法。其复杂重建仅需数分钟,远低于现有方法;训练速度达最先进基于NeRF模型的110倍;在相同运行时间限制下,性能显著优于现有动态人像NeRF方法;仅需30秒训练即可提供有效视觉输出。
原文摘要 · Abstract (English)
High-fidelity digital human representations are increasingly in demand in the digital world, particularly for interactive telepresence, AR/VR, 3D graphics, and the rapidly evolving metaverse. Even though they work well in small spaces, conventional methods for reconstructing 3D human motion frequently require the use of expensive hardware and have high processing costs. This study presents HumanAvatar, an innovative approach that efficiently reconstructs precise human avatars from monocular video sources. At the core of our methodology, we integrate the pre-trained HuMoR, a model celebrated for its proficiency in human motion estimation. This is adeptly fused with the cutting-edge neural radiance field technology, Instant-NGP, and the state-of-the-art articulated model, Fast-SNARF, to enhance the reconstruction fidelity and speed. By combining these two technologies, a system is created that can render quickly and effectively while also providing estimation of human pose parameters that are unmatched in accuracy. We have enhanced our system with an advanced posture-sensitive space reduction technique, which optimally balances rendering quality with computational efficiency. In our detailed experimental analysis using both artificial and real-world monocular videos, we establish the advanced performance of our approach. HumanAvatar consistently equals or surpasses contemporary leading-edge reconstruction techniques in quality. Furthermore, it achieves these complex reconstructions in minutes, a fraction of the time typically required by existing methods. Our models achieve a training speed that is 110X faster than that of State-of-The-Art (SoTA) NeRF-based models. Our technique performs noticeably better than SoTA dynamic human NeRF methods if given an identical runtime limit. HumanAvatar can provide effective visuals after only 30 seconds of training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。