从被遮挡的2D图像重建完整3D物体,突破传统方法局限。
Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images
- 基于掩码加权注意力与遮挡感知机制,显式利用遮挡先验
- 仅用合成数据训练,在真实场景中实现完整3D重建
- 比先完成2D补全再做3D重建的方法更优,适合真实世界应用
大多数基于图像的3D重建方法假设物体完全可见,忽视了现实场景中常见的遮挡问题。本文提出Amodal3R,一种条件3D生成模型,旨在从部分观测中重建3D物体。该模型以基础3D生成模型为基础,引入掩码加权多头交叉注意力机制和遮挡感知注意力层,显式利用遮挡先验指导重建过程。实验表明,仅在合成数据上训练的Amodal3R,能在真实场景的遮挡条件下恢复出完整的3D几何与外观。其性能显著优于独立进行2D非遮挡补全后进行3D重建的现有方法,为遮挡感知3D重建设立了新基准。
原文摘要 · Abstract (English)
Most image-based 3D object reconstructors assume that objects are fully visible, ignoring occlusions that commonly occur in real-world scenarios. In this paper, we introduce Amodal3R, a conditional 3D generative model designed to reconstruct 3D objects from partial observations. We start from a "foundation" 3D generative model and extend it to recover plausible 3D geometry and appearance from occluded objects. We introduce a mask-weighted multi-head cross-attention mechanism followed by an occlusion-aware attention layer that explicitly leverages occlusion priors to guide the reconstruction process. We demonstrate that, by training solely on synthetic data, Amodal3R learns to recover full 3D objects even in the presence of occlusions in real scenes. It substantially outperforms existing methods that independently perform 2D amodal completion followed by 3D reconstruction, thereby establishing a new benchmark for occlusion-aware 3D reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。