arXiv:2506.21544cs.CV2025-06被引 12

从单张遮挡图像生成六视角3D结构,无需预处理或标注

DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View Diffusion

  • 自监督学习利用遮挡-未遮挡图像对训练模型补全结构
  • 直接生成六幅结构一致的新视角,重建质量提升32%以上
  • 首个面向遮挡重建的基准测试,覆盖多种遮挡场景

从单张图像重建3D物体仍具挑战性,尤其在真实遮挡情况下。尽管近期基于扩散的视图合成模型可从单张RGB图像生成一致的新视角,但通常假设输入完全可见,当物体部分被遮挡时,3D重建质量显著下降。我们提出DeOcc-1-to-3,一个端到端的遮挡感知多视角生成框架,可直接从单张遮挡图像生成六幅结构一致的新视角,实现可靠的3D重建,无需前期修补或人工标注。自监督训练流程利用遮挡-未遮挡图像对和伪真值视图,指导模型学习结构感知补全与视图一致性。在不修改原始架构的前提下,对视图合成模型进行全微调,联合学习补全与多视角生成。此外,我们构建了首个面向遮挡感知重建的基准测试,涵盖多样化的遮挡程度、物体类别和掩码模式,为未来评估提供标准化协议。

原文摘要 · Abstract (English)

Reconstructing 3D objects from a single image remains challenging, especially under real-world occlusions. While recent diffusion-based view synthesis models can generate consistent novel views from a single RGB image, they typically assume fully visible inputs and fail when parts of the object are occluded, resulting in degraded 3D reconstruction quality. We propose DeOcc-1-to-3, an end-to-end framework for occlusion-aware multi-view generation that synthesizes six structurally consistent novel views directly from a single occluded image, enabling reliable 3D reconstruction without prior inpainting or manual annotations. Our self-supervised training pipeline leverages occluded-unoccluded image pairs and pseudo-ground-truth views to teach the model structure-aware completion and view consistency. Without modifying the original architecture, we fully fine-tune the view synthesis model to jointly learn completion and multi-view generation. Additionally, we introduce the first benchmark for occlusion-aware reconstruction, covering diverse occlusion levels, object categories, and masking patterns, providing a standardized protocol for future evaluation.

3D重建扩散模型遮挡补全自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。