仅用一张照片生成可动画3D人像,支持不同姿态与视角
Bringing Your Portrait to 3D Presence
- 采用双UV表示法消除姿态和构图带来的特征偏差
- 仅用半身合成数据训练,实现头部与上半身顶尖重建效果
- 适合虚拟形象、数字人开发等需要快速建模的场景
我们提出一个统一框架,仅需单张头像、半身或全身照片即可重建可动画的3D人体模型。针对姿态与构图敏感的特征表示、可扩展数据有限、代理网格估计不可靠三大瓶颈,提出双UV表示法(Core-UV与Shell-UV分支),将图像特征映射至标准UV空间,消除姿态与构图引起的特征偏移;构建结合2D生成多样性与几何一致3D渲染的因子化合成数据集,并设计训练方案提升真实感与身份一致性;引入鲁棒代理网格追踪器,在部分遮挡下仍保持稳定。仅在半身合成数据上训练,模型在头部与上半身重建上达到当前最佳表现,全身体积重建结果也具竞争力。大量实验与分析验证了方法有效性。
原文摘要 · Abstract (English)
We present a unified framework for reconstructing animatable 3D human avatars from a single portrait across head, half-body, and full-body inputs. Our method tackles three bottlenecks: pose- and framing-sensitive feature representations, limited scalable data, and unreliable proxy-mesh estimation. We introduce a Dual-UV representation that maps image features to a canonical UV space via Core-UV and Shell-UV branches, eliminating pose- and framing-induced token shifts. We also build a factorized synthetic data manifold combining 2D generative diversity with geometry-consistent 3D renderings, supported by a training scheme that improves realism and identity consistency. A robust proxy-mesh tracker maintains stability under partial visibility. Together, these components enable strong in-the-wild generalization. Trained only on half-body synthetic data, our model achieves state-of-the-art head and upper-body reconstruction and competitive full-body results. Extensive experiments and analyses further validate the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。