用预训练感知模型提升单图3D重建质量
Seeing Before Generating: Object Perception Enhances Single-View 3D Reconstruction

- 利用预训练模型提取语义与几何感知信号驱动重建
- 在基准数据集上显著提升两种先进重建方法性能
- 可插即用,适用于多种3D重建框架
人类视觉中对象感知与重建的关系已明确,但在计算机视觉中仍研究不足。本文证明,学习到的对象感知能显著提升3D重建效果。针对单视图3D物体重建这一挑战性任务,我们提出一种方法:利用从预训练感知模型中提取的语义和几何信息作为感知信号,指导单张图像的物体重建。该方法具备模型无关性,可无缝集成到多种重建流程中,实现即插即用。在基准数据集上的实验表明,该方法对两种顶尖的单视图3D重建流水线均带来一致且显著的性能提升,验证了将感知融入生成的有效性。本文还对方法的多个方面进行了深入分析。
原文摘要 · Abstract (English)
The relationship between object perception and reconstruction is well established in human vision, yet remains underexplored in computer vision. In this paper, we demonstrate that learnt object perception can significantly enhance 3D reconstruction. Focusing on the challenging task of single-view 3D object reconstruction, we propose a method that leverages perceptual signals extracted from pretrained perception models capturing semantic and geometric information to drive the reconstruction of an object from its single image. Our approach is model-agnostic and can be integrated into various reconstruction methods in a plug-and-play manner. Experiments with two state-of-the-art single-view 3D reconstruction pipelines in a benchmark dataset show consistent and substantial improvements achieved by our method, validating the effectiveness of incorporating perception into generation. We provide in-depth analysis of various aspects of our method and its application. Our project page is at https://ynhuhuynh.github.io/perception-3d/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。