仅用几张图生成可动3D虚拟人,不依赖姿态信息
NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human Poses
- 纯图像输入重建,无需人体姿态数据
- 在无真实姿态条件下性能超越现有方法
- 适合实际场景应用,对姿态误差不敏感
本文研究从单张或稀疏图像中恢复可动画化的3D人体模型。现有先进方法通常依赖精确的“真值”相机位姿和人体姿态来指导重建,但实验表明,当姿态估计存在噪声时,依赖姿态的重建性能显著下降。为此,我们提出NoPo-Avatar,仅通过图像输入实现人体重建,完全不依赖测试时的人体姿态。该方法摆脱了对姿态的依赖,因此不受姿态估计误差影响,适用性更广。在THuman2.0、XHuman和HuGe100K等挑战性数据集上的实验显示,NoPo-Avatar在无真值姿态的实际场景中优于现有基线,在有真值姿态的实验室环境下也达到可比性能。
原文摘要 · Abstract (English)
We tackle the task of recovering an animatable 3D human avatar from a single or a sparse set of images. For this task, beyond a set of images, many prior state-of-the-art methods use accurate "ground-truth" camera poses and human poses as input to guide reconstruction at test-time. We show that pose-dependent reconstruction degrades results significantly if pose estimates are noisy. To overcome this, we introduce NoPo-Avatar, which reconstructs avatars solely from images, without any pose input. By removing the dependence of test-time reconstruction on human poses, NoPo-Avatar is not affected by noisy human pose estimates, making it more widely applicable. Experiments on challenging THuman2.0, XHuman, and HuGe100K data show that NoPo-Avatar outperforms existing baselines in practical settings (without ground-truth poses) and delivers comparable results in lab settings (with ground-truth poses).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。