用日常穿搭照快速重建高精度3D虚拟形象,支持动作与遮挡处理。
PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day Photos
- 直接建模全身外观,避免分割误差,结合姿态控制与新损失函数优化细节。
- 5分钟完成个性化重建,速度比之前方法快48倍,保留头发等高频纹理。
- 基于神经辐射场的连续表示,正确处理遮挡,适合虚拟试穿与动画应用。
我们提出PFAvatar(Pose-Fusion Avatar),一种从真实世界穿搭照片(OOTD)中重建高质量3D虚拟形象的新方法,该方法可处理多姿态、遮挡和复杂背景。流程分为两阶段:(1) 利用少量OOTD样本微调一个姿态感知的扩散模型;(2) 将3D虚拟形象以神经辐射场(NeRF)形式蒸馏。第一阶段不进行服装部件分割,而是直接建模全身外观,通过预训练的ControlNet实现姿态估计,并引入新型条件先验保持损失(CPPL),在少样本训练中实现端到端细节学习并缓解语言漂移。整个个性化过程仅需5分钟,相比以往方法提速48倍。第二阶段采用基于标准SMPL-X空间采样与多分辨率3D-SDS优化的NeRF表示,相较于网格类方法存在的离散化局限与遮挡几何错误,其连续辐射场能有效保留高频纹理(如头发),并通过透射率正确处理遮挡。实验表明,PFAvatar在重建保真度、细节保留和对遮挡/截断的鲁棒性方面优于现有最优方法,推动了真实世界OOTD图集生成实用3D虚拟形象的发展。此外,重建的3D虚拟形象可支持虚拟试穿、动画与人物视频重演等下游应用,展现其多样性和实用性。
原文摘要 · Abstract (English)
We propose PFAvatar (Pose-Fusion Avatar), a new method that reconstructs high-quality 3D avatars from Outfit of the Day(OOTD) photos, which exhibit diverse poses, occlusions, and complex backgrounds. Our method consists of two stages: (1) fine-tuning a pose-aware diffusion model from few-shot OOTD examples and (2) distilling a 3D avatar represented by a neural radiance field (NeRF). In the first stage, unlike previous methods that segment images into assets (e.g., garments, accessories) for 3D assembly, which is prone to inconsistency, we avoid decomposition and directly model the full-body appearance. By integrating a pre-trained ControlNet for pose estimation and a novel Condition Prior Preservation Loss (CPPL), our method enables end-to-end learning of fine details while mitigating language drift in few-shot training. Our method completes personalization in just 5 minutes, achieving a 48x speed-up compared to previous approaches. In the second stage, we introduce a NeRF-based avatar representation optimized by canonical SMPL-X space sampling and Multi-Resolution 3D-SDS. Compared to mesh-based representations that suffer from resolution-dependent discretization and erroneous occluded geometry, our continuous radiance field can preserve high-frequency textures (e.g., hair) and handle occlusions correctly through transmittance. Experiments demonstrate that PFAvatar outperforms state-of-the-art methods in terms of reconstruction fidelity, detail preservation, and robustness to occlusions/truncations, advancing practical 3D avatar generation from real-world OOTD albums. In addition, the reconstructed 3D avatar supports downstream applications such as virtual try-on, animation, and human video reenactment, further demonstrating the versatility and practical value of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。