arXiv:2503.19207cs.CV2025-03CVPR被引 7

仅用几张照片即可快速生成逼真可动的3D人物模型,无需繁琐优化。

FRESA: Feedforward Reconstruction of Personalized Skinned Avatars from Few Images

  • 联合推断个性化形状、权重和姿态变形,提升几何精度
  • 多帧特征聚合减少伪影,保留人物身份特征
  • 训练后直接处理手机拍摄照片,零样本泛化能力强

我们提出一种新方法,仅需少量图像即可重建个性化的3D人体角色并实现真实动画。由于体型、姿势和服装差异大,现有方法在推理时大多需要数小时的个体优化,限制了实际应用。本文通过学习上千名穿衣服人的通用先验,实现即时前向生成和零样本泛化。具体而言,不使用共享蒙皮权重,而是联合推断个性化角色形状、蒙皮权重和姿态依赖形变,有效提升整体几何保真度并减少形变伪影。为消除姿态变化带来的模糊性,设计3D规范化流程以生成像素对齐的初始条件,有助于恢复精细几何细节。随后采用多帧特征聚合方法,鲁棒地降低规范化引入的伪影,并融合出保留个体特征的合理角色。模型在大规模捕获数据集上端到端训练,包含多样化人体与高质量3D扫描配对数据。大量实验表明,本方法生成结果比当前最优方法更逼真,且可直接应用于随意拍摄的手机照片。

原文摘要 · Abstract (English)

We present a novel method for reconstructing personalized 3D human avatars with realistic animation from only a few images. Due to the large variations in body shapes, poses, and cloth types, existing methods mostly require hours of per-subject optimization during inference, which limits their practical applications. In contrast, we learn a universal prior from over a thousand clothed humans to achieve instant feedforward generation and zero-shot generalization. Specifically, instead of rigging the avatar with shared skinning weights, we jointly infer personalized avatar shape, skinning weights, and pose-dependent deformations, which effectively improves overall geometric fidelity and reduces deformation artifacts. Moreover, to normalize pose variations and resolve coupled ambiguity between canonical shapes and skinning weights, we design a 3D canonicalization process to produce pixel-aligned initial conditions, which helps to reconstruct fine-grained geometric details. We then propose a multi-frame feature aggregation to robustly reduce artifacts introduced in canonicalization and fuse a plausible avatar preserving person-specific identities. Finally, we train the model in an end-to-end framework on a large-scale capture dataset, which contains diverse human subjects paired with high-quality 3D scans. Extensive experiments show that our method generates more authentic reconstruction and animation than state-of-the-arts, and can be directly generalized to inputs from casually taken phone photos. Project page and code is available at https://github.com/rongakowang/FRESA.

3D重建人物建模即时生成少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。