arXiv:2503.10625cs.CVcs.AI2025-03被引 66

单图秒级生成可动画3D人体,精度与泛化能力双突破

LHM: Large Animatable Human Reconstruction Model from a Single Image in Seconds

  • 用多模态变压器+高斯点云,端到端重建人体几何与纹理
  • 生成速度仅需数秒,无需后处理,面部与手部细节完整
  • 适合虚拟人、游戏开发等需要快速建模的场景

从单张图像实现可动画3D人体重建极具挑战性,因需解耦几何、外观与形变。现有方法多聚焦静态建模,依赖合成3D扫描训练,泛化能力受限;基于优化的视频方法虽精度高,但需受控拍摄条件且计算开销大。受大型静态重建模型启发,我们提出LHM(Large Animatable Human Reconstruction Model),可在前馈过程中以3D高斯点云形式生成高保真虚拟形象。模型采用多模态变换器架构,通过注意力机制融合人体位置特征与图像特征,有效保留衣物几何与纹理细节。为增强面部身份一致性与微细节恢复,设计头部特征金字塔编码方案,聚合多尺度头部特征。大量实验表明,LHM可在数秒内生成无后处理的可动画人体,且在重建精度与泛化能力上均优于现有方法。

原文摘要 · Abstract (English)

Animatable 3D human reconstruction from a single image is a challenging problem due to the ambiguity in decoupling geometry, appearance, and deformation. Recent advances in 3D human reconstruction mainly focus on static human modeling, and the reliance of using synthetic 3D scans for training limits their generalization ability. Conversely, optimization-based video methods achieve higher fidelity but demand controlled capture conditions and computationally intensive refinement processes. Motivated by the emergence of large reconstruction models for efficient static reconstruction, we propose LHM (Large Animatable Human Reconstruction Model) to infer high-fidelity avatars represented as 3D Gaussian splatting in a feed-forward pass. Our model leverages a multimodal transformer architecture to effectively encode the human body positional features and image features with attention mechanism, enabling detailed preservation of clothing geometry and texture. To further boost the face identity preservation and fine detail recovery, we propose a head feature pyramid encoding scheme to aggregate multi-scale features of the head regions. Extensive experiments demonstrate that our LHM generates plausible animatable human in seconds without post-processing for face and hands, outperforming existing methods in both reconstruction accuracy and generalization ability.

3D重建单图生成可动画模型高斯点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。