用Transformer实现单目或稀疏图像下的人体重建与可控动画
HumanRAM: Feed-forward Human Reconstruction and Animation Model using Transformers
- 通过共享SMPL-X神经纹理和显式姿态条件,统一建模人体重建与动画
- 在真实数据集上重建精度、动画保真度和泛化性能显著优于已有方法
- 支持单目或稀疏输入,无需逐人优化,适合快速生成高质量人体视频
3D人体重建与动画是计算机图形学与视觉领域的长期课题。现有方法通常依赖复杂的多视角采集和耗时的逐主体优化。为此,我们提出HumanRAM,一种新颖的前馈式通用人体重建与动画方法,仅需单目或稀疏人体图像即可完成。该方法将人体重建与动画统一建模,通过引入由共享SMPL-X神经纹理参数化的显式姿态条件,嵌入基于Transformer的大规模重建模型(LRM)。给定带有相机参数和SMPL-X姿态的单目或稀疏输入图像,模型利用可扩展的Transformer和基于DPT的解码器,合成新视角和新姿态下的逼真人像渲染结果。通过显式姿态条件,模型同时实现高精度人体重建与高保真度的姿态控制动画。实验表明,HumanRAM在真实世界数据集上的重建准确率、动画保真度和泛化能力均显著超越此前方法。视频演示见 https://zju3dv.github.io/humanram/。
原文摘要 · Abstract (English)
3D human reconstruction and animation are long-standing topics in computer graphics and vision. However, existing methods typically rely on sophisticated dense-view capture and/or time-consuming per-subject optimization procedures. To address these limitations, we propose HumanRAM, a novel feed-forward approach for generalizable human reconstruction and animation from monocular or sparse human images. Our approach integrates human reconstruction and animation into a unified framework by introducing explicit pose conditions, parameterized by a shared SMPL-X neural texture, into transformer-based large reconstruction models (LRM). Given monocular or sparse input images with associated camera parameters and SMPL-X poses, our model employs scalable transformers and a DPT-based decoder to synthesize realistic human renderings under novel viewpoints and novel poses. By leveraging the explicit pose conditions, our model simultaneously enables high-quality human reconstruction and high-fidelity pose-controlled animation. Experiments show that HumanRAM significantly surpasses previous methods in terms of reconstruction accuracy, animation fidelity, and generalization performance on real-world datasets. Video results are available at https://zju3dv.github.io/humanram/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。