仅用一张照片生成1024分辨率的全身多视角视频。
Pippo: High-Resolution Multi-View Humans from a Single Image
- 基于扩散变换器,无需参数化人体模型或相机参数。
- 训练中低分辨率多视角去噪,推理时生成超5倍于训练视角数的图像。
- 通过空间锚点与普鲁克射线实现3D一致性生成,适合数字人建模应用。
我们提出Pippo,一种生成模型,可从单张随意拍摄的照片生成1024分辨率的稠密全身多视角视频。Pippo为多视角扩散变换器,不需额外输入(如拟合参数化人体模型或图像相机参数)。在无标注的30亿张人体图像上预训练,再对工作室拍摄的人体进行多视角中段和后段训练。中段训练时,在低分辨率下对多达48个视角进行去噪,并用浅层MLP粗略编码目标相机。后段训练时,在高分辨率下对较少视角进行去噪,并使用像素对齐控制(如空间锚点和普鲁克射线)以实现3D一致生成。推理阶段,我们提出注意力偏置技术,使Pippo能同时生成超过训练时所见视角数5倍的图像。此外,我们引入一种改进的评估指标来衡量多视角生成的3D一致性,结果显示Pippo在单图生成多视角人体方面优于现有方法。
原文摘要 · Abstract (English)
We present Pippo, a generative model capable of producing 1K resolution dense turnaround videos of a person from a single casually clicked photo. Pippo is a multi-view diffusion transformer and does not require any additional inputs - e.g., a fitted parametric model or camera parameters of the input image. We pre-train Pippo on 3B human images without captions, and conduct multi-view mid-training and post-training on studio captured humans. During mid-training, to quickly absorb the studio dataset, we denoise several (up to 48) views at low-resolution, and encode target cameras coarsely using a shallow MLP. During post-training, we denoise fewer views at high-resolution and use pixel-aligned controls (e.g., Spatial anchor and Plucker rays) to enable 3D consistent generations. At inference, we propose an attention biasing technique that allows Pippo to simultaneously generate greater than 5 times as many views as seen during training. Finally, we also introduce an improved metric to evaluate 3D consistency of multi-view generations, and show that Pippo outperforms existing works on multi-view human generation from a single image.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。