单张照片生成可动画的高细节3D人像,适合真实应用。
AdaHuman: Animatable Detailed 3D Human Generation with Compositional Multiview Diffusion

- 用多视角扩散模型生成带姿态控制的3D高斯点云。
- 局部细节通过图像到图像精修提升,减少自遮挡。
- 生成结果支持任意动作驱动,适合游戏与影视制作。
现有图像转3D人像方法难以生成适用于实际场景的高细节、可动画化人像。本文提出AdaHuman框架,仅需一张自然场景图像即可生成高保真可动画3D人像。核心创新包括:(1) 姿态条件化的3D关节扩散模型,在每一步扩散过程中同步生成任意姿态下的多视角一致图像及对应的3D高斯点云(3DGS);(2) 组合式3DGS精修模块,通过图像到图像精修增强局部身体部位细节,并利用新型裁剪感知相机射线图无缝融合,生成连贯的高细节3D人像。该方法可生成高度逼真的标准A姿态人像,自遮挡极小,支持任意输入动作的绑定与动画。在公开基准和自然图像上的大量评估表明,AdaHuman在人像重建与重姿态表现上显著优于当前最先进方法。代码与模型将开源供研究使用。
原文摘要 · Abstract (English)
Existing methods for image-to-3D avatar generation struggle to produce highly detailed, animation-ready avatars suitable for real-world applications. We introduce AdaHuman, a novel framework that generates high-fidelity animatable 3D avatars from a single in-the-wild image. AdaHuman incorporates two key innovations: (1) A pose-conditioned 3D joint diffusion model that synthesizes consistent multi-view images in arbitrary poses alongside corresponding 3D Gaussian Splats (3DGS) reconstruction at each diffusion step; (2) A compositional 3DGS refinement module that enhances the details of local body parts through image-to-image refinement and seamlessly integrates them using a novel crop-aware camera ray map, producing a cohesive detailed 3D avatar. These components allow AdaHuman to generate highly realistic standardized A-pose avatars with minimal self-occlusion, enabling rigging and animation with any input motion. Extensive evaluation on public benchmarks and in-the-wild images demonstrates that AdaHuman significantly outperforms state-of-the-art methods in both avatar reconstruction and reposing. Code and models will be publicly available for research purposes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。