arXiv:2606.02573cs.CV2026-06

单图生成逼真3D人像,1秒完成且无需优化

HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single Image

论文配图:HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single Image
图 1 · 摘自论文原文
  • 用两种数据策略构建10万+高质量3D人像数据集
  • 输入单张照片和简化人体网格,1秒生成3D人像
  • 适合游戏、影视等需要快速生成高保真人像的场景

本文提出HumanNOVA,一种从单张RGB图像快速生成逼真、通用3D人像的模型。由于高质量多样化的3D人像数据稀缺,我们设计可扩展的数据生成管道:一是利用现有绑定资产并注入日常生活中的丰富姿态;二是基于多相机捕获的人体数据,通过拟合生成更多视角用于训练。该方法共生成10万+资产,显著提升训练数据的数量与多样性。在架构上,HumanNOVA采用前馈式、令牌条件化的建模框架,无需测试时优化,推理时间小于1秒。给定输入图像和估计的简化人体网格(SMPL),模型先将两者编码为紧凑令牌表示,再通过交叉注意力融合,构建基于三平面的3D人像表示。在多个基准上的实验表明,该方法在定量与定性指标上均表现优越,且对不同输入图像具有强鲁棒性。

原文摘要 · Abstract (English)

In this paper, we present HumanNOVA, a photorealistic, universal, and rapid model for generating 3D human avatars from a single RGB image. Achieving both photorealism and generalization is challenging due to the scarcity of diverse, high-quality 3D human data. To address this, we build a scalable data generation pipeline that follows two strategies. The first one is to leverage existing rigged assets and animate them with extensive poses from daily life. The second strategy is to utilize existing multi-camera captures of humans and employ fitting to generate more diverse views for training. These two strategies enable us to scale up to 100k assets, significantly enhancing both the quantity and the diversity of data for robust model training. In terms of the architecture, HumanNOVA adopts a feed-forward, token-conditioned avatar modeling framework that allows fast inference in less than one second and requires no test-time optimization. Given an input image and an estimated simplified human mesh (SMPL) without detailed geometry or appearance, the model first encodes both inputs into compact token representations. These tokens then act as conditioning signals and are fused through cross-attention to construct a triplane-based 3D avatar representation. Extensive experiments on multiple benchmarks demonstrate the superiority of our approach, both quantitatively and qualitatively, as well as its robustness under diverse input image conditions. Project page at https://HumanNOVA.github.io .

3D人像单图生成快速建模逼真渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。