arXiv:2506.13766cs.CV2025-06被引 5

无需姿态或相机信息,秒级生成可动画3D人体模型

LHM++: An Efficient Large Human Reconstruction Model for Pose-free Images to 3D

  • 用点图Transformer融合图像与3D点特征,提升重建效率
  • 生成高质量3D高斯点云,支持实时动画渲染
  • 适合快速建模、虚拟人开发等应用

从无姿态、无相机参数的随意拍摄图像中重建可动画3D人体,具有高度实用性但面临视角错位、遮挡和缺乏结构先验等挑战。本文提出LHM++,一个高效的大规模人体重建模型,仅需数秒即可从单张或多张姿态无关图像生成高质量可动画3D化身。核心采用编码器-解码器点图变换架构,逐步编码并解码3D几何点特征,通过多模态注意力融合层次化3D点特征与图像特征,并解码为3D高斯点以恢复精细几何与外观。为进一步提升视觉保真度,引入轻量级3D感知神经动画渲染器,实现实时渲染质量优化。大量实验表明,该方法无需相机或姿态标注即可生成高保真、可动画3D人体。代码与项目页见https://lingtengqiu.github.io/LHM++/

原文摘要 · Abstract (English)

Reconstructing animatable 3D humans from casually captured images of articulated subjects without camera or pose information is highly practical but remains challenging due to view misalignment, occlusions, and the absence of structural priors. In this work, we present LHM++, an efficient large-scale human reconstruction model that generates high-quality, animatable 3D avatars within seconds from one or multiple pose-free images. At its core is an Encoder-Decoder Point-Image Transformer architecture that progressively encodes and decodes 3D geometric point features to improve efficiency, while fusing hierarchical 3D point features with image features through multimodal attention. The fused features are decoded into 3D Gaussian splats to recover detailed geometry and appearance. To further enhance visual fidelity, we introduce a lightweight 3D-aware neural animation renderer that refines the rendering quality of reconstructed avatars in real time. Extensive experiments show that our method produces high-fidelity, animatable 3D humans without requiring camera or pose annotations. Our code and project page are available at https://lingtengqiu.github.io/LHM++/

3D重建神经渲染虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。