单图重建穿衣服的人体,还能抗遮挡、多视角一致。
CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-Consistency from a Single Image
- 用多视角扩散模型生成无遮挡图像,保持视角一致性。
- 无需真实3D标注,在遮挡情况下仍能提升3dB PSNR。
- 适合野外复杂场景下人体3D重建,无需昂贵标注。
从单张图像重建穿衣服的人体是计算机视觉中的基础任务,应用广泛。现有方法通常假设人体处于无遮挡环境,面对真实世界中的遮挡图像时,会生成多视角不一致且碎片化的结果。此外,多数方法依赖如SMPL等几何先验进行训练与推理,而这些标注在实际中难以获取。为此,我们提出CHROME:一种从单张遮挡图像中重建抗遮挡、多视角一致的3D人体的新方法,无需真实几何先验或3D监督。CHROME首先利用多视角扩散模型从遮挡输入中合成无遮挡人体图像,并支持现成的姿态控制以显式保证跨视角一致性。随后,一个3D重建模型基于遮挡输入和合成视图预测一组3D高斯点云,对齐跨视角细节,生成连贯准确的3D表示。CHROME在新视角合成(最高提升3 dB PSNR)和复杂条件下的几何重建方面均取得显著进步。
原文摘要 · Abstract (English)
Reconstructing clothed humans from a single image is a fundamental task in computer vision with wide-ranging applications. Although existing monocular clothed human reconstruction solutions have shown promising results, they often rely on the assumption that the human subject is in an occlusion-free environment. Thus, when encountering in-the-wild occluded images, these algorithms produce multiview inconsistent and fragmented reconstructions. Additionally, most algorithms for monocular 3D human reconstruction leverage geometric priors such as SMPL annotations for training and inference, which are extremely challenging to acquire in real-world applications. To address these limitations, we propose CHROME: Clothed Human Reconstruction with Occlusion-Resilience and Multiview-ConsistEncy from a Single Image, a novel pipeline designed to reconstruct occlusion-resilient 3D humans with multiview consistency from a single occluded image, without requiring either ground-truth geometric prior annotations or 3D supervision. Specifically, CHROME leverages a multiview diffusion model to first synthesize occlusion-free human images from the occluded input, compatible with off-the-shelf pose control to explicitly enforce cross-view consistency during synthesis. A 3D reconstruction model is then trained to predict a set of 3D Gaussians conditioned on both the occluded input and synthesized views, aligning cross-view details to produce a cohesive and accurate 3D representation. CHROME achieves significant improvements in terms of both novel view synthesis (upto 3 db PSNR) and geometric reconstruction under challenging conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。