从单张图重建带衣服的人和物,细节更真实。
Realistic Clothed Human and Object Joint Reconstruction from a Single Image
- 用隐式表示联合建模人与物体,捕捉衣物等细节。
- 引入注意力机制与扩散修复,解决遮挡导致的细节丢失。
- 适合需要高保真人物与物体三维重建的研究者。
近期的单图像联合重建方法多采用基于模板或粗略模型表示3D形状,难以捕捉人体上松散衣物的细节。本文提出一种新型隐式方法,首次在单目视角下联合重建真实感3D着装人体与物体。该任务因人体-物体相互遮挡及2D图像缺乏3D信息而极具挑战,常导致细节重建差和深度模糊。为此,我们设计了一种基于注意力的神经隐式模型,利用输入图像中人体-物体整体像素对齐进行全局理解,并结合人体与物体的局部独立视图提升现实感(如衣物细节)。同时,网络以估计的人体-物体姿态先验提取的语义特征为条件,提供两者共享空间的3D位置信息。为处理物体遮挡带来的人体缺失区域,采用生成式扩散模型进行补全,恢复丢失细节。我们构建了一个合成数据集,包含渲染的相互遮挡3D人体扫描与多样化物体场景。在合成与真实世界数据集上的大量实验表明,所提方法在重建质量上优于现有竞争方法。
原文摘要 · Abstract (English)
Recent approaches to jointly reconstruct 3D humans and objects from a single RGB image represent 3D shapes with template-based or coarse models, which fail to capture details of loose clothing on human bodies. In this paper, we introduce a novel implicit approach for jointly reconstructing realistic 3D clothed humans and objects from a monocular view. For the first time, we model both the human and the object with an implicit representation, allowing to capture more realistic details such as clothing. This task is extremely challenging due to human-object occlusions and the lack of 3D information in 2D images, often leading to poor detail reconstruction and depth ambiguity. To address these problems, we propose a novel attention-based neural implicit model that leverages image pixel alignment from both the input human-object image for a global understanding of the human-object scene and from local separate views of the human and object images to improve realism with, for example, clothing details. Additionally, the network is conditioned on semantic features derived from an estimated human-object pose prior, which provides 3D spatial information about the shared space of humans and objects. To handle human occlusion caused by objects, we use a generative diffusion model that inpaints the occluded regions, recovering otherwise lost details. For training and evaluation, we introduce a synthetic dataset featuring rendered scenes of inter-occluded 3D human scans and diverse objects. Extensive evaluation on both synthetic and real-world datasets demonstrates the superior quality of the proposed human-object reconstructions over competitive methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。