无需调参,一键从乱拍照片重建高精度穿衣3D人像。
UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections
- 直接处理随意拍摄的杂乱照片,秒级生成标准多视角图像。
- 几何精度提升15%-18%,纹理保真度提升21%-46%。
- 支持任意姿态控制,适合日常照片的3D虚拟试衣场景。
我们提出UP2You,首个无需调参的方案,可从极不约束的野外2D照片中重建高保真穿衣3D人像。不同于以往需干净输入(如无遮挡全身照或校准视图)的方法,UP2You直接处理原始、无序的照片,这些照片可能在姿态、视角、裁剪和遮挡上差异极大。我们引入数据矫正范式,在单次前向传播中数秒内将无约束输入转化为清晰、正交的多视角图像,简化3D重建流程。核心是姿态相关特征聚合模块(PCFA),按目标姿态选择性融合多参考图像信息,实现更好身份保留且内存恒定。还设计基于Perceiver的多参考形状预测器,无需预捕获人体模板。在4D-Dress、PuzzleIOI及野外采集数据上实验表明,UP2You在几何精度(PuzzleIOI上Chamfer-15%,P2S-18%)与纹理保真度(4D-Dress上PSNR-21%,LPIPS-46%)上持续领先。整个流程仅需1.5分钟/人,具备任意姿态控制与零训练多服装3D虚拟试衣能力,适用于真实场景。模型与代码将公开,推动该未充分研究任务发展。
原文摘要 · Abstract (English)
We present UP2You, the first tuning-free solution for reconstructing high-fidelity 3D clothed portraits from extremely unconstrained in-the-wild 2D photos. Unlike previous approaches that require "clean" inputs (e.g., full-body images with minimal occlusions, or well-calibrated cross-view captures), UP2You directly processes raw, unstructured photographs, which may vary significantly in pose, viewpoint, cropping, and occlusion. Instead of compressing data into tokens for slow online text-to-3D optimization, we introduce a data rectifier paradigm that efficiently converts unconstrained inputs into clean, orthogonal multi-view images in a single forward pass within seconds, simplifying the 3D reconstruction. Central to UP2You is a pose-correlated feature aggregation module (PCFA), that selectively fuses information from multiple reference images w.r.t. target poses, enabling better identity preservation and nearly constant memory footprint, with more observations. We also introduce a perceiver-based multi-reference shape predictor, removing the need for pre-captured body templates. Extensive experiments on 4D-Dress, PuzzleIOI, and in-the-wild captures demonstrate that UP2You consistently surpasses previous methods in both geometric accuracy (Chamfer-15%, P2S-18% on PuzzleIOI) and texture fidelity (PSNR-21%, LPIPS-46% on 4D-Dress). UP2You is efficient (1.5 minutes per person), and versatile (supports arbitrary pose control, and training-free multi-garment 3D virtual try-on), making it practical for real-world scenarios where humans are casually captured. Both models and code will be released to facilitate future research on this underexplored task. Project Page: https://zcai0612.github.io/UP2You
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。