用先验知识提升3D场景坐标重建的准确性与稳定性。
Scene Coordinate Reconstruction Priors
- 引入概率视角,融合高阶几何先验优化场景坐标回归
- 在三个室内数据集上提升点云一致性与相机定位成功率
- 适合需要鲁棒3D重建的视觉定位与新视角合成任务
场景坐标回归(SCR)模型在3D视觉中表现出强大的隐式场景表征能力,可用于视觉重定位和运动结构重建。然而,当训练图像提供的多视角约束不足时,模型性能会退化。本文提出对训练过程的概率重解释,使我们能够注入高层重建先验。研究了多种先验,从深度值分布的简单先验到对合理场景坐标配置的可学习先验。后者通过在大量室内扫描数据上训练3D点云扩散模型获得。这些先验在每一步训练中引导预测的3D点向更合理的几何结构靠拢,从而提高其可能性。在三个室内数据集上,该方法显著提升了场景表征质量,生成更连贯的点云、更高的配准率和更优的相机位姿,对新视角合成与相机重定位等下游任务均有积极影响。
原文摘要 · Abstract (English)
Scene coordinate regression (SCR) models have proven to be powerful implicit scene representations for 3D vision, enabling visual relocalization and structure-from-motion. SCR models are trained specifically for one scene. If training images imply insufficient multi-view constraints SCR models degenerate. We present a probabilistic reinterpretation of training SCR models, which allows us to infuse high-level reconstruction priors. We investigate multiple such priors, ranging from simple priors over the distribution of reconstructed depth values to learned priors over plausible scene coordinate configurations. For the latter, we train a 3D point cloud diffusion model on a large corpus of indoor scans. Our priors push predicted 3D scene points towards plausible geometry at each training step to increase their likelihood. On three indoor datasets our priors help learning better scene representations, resulting in more coherent scene point clouds, higher registration rates and better camera poses, with a positive effect on down-stream tasks such as novel view synthesis and camera relocalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。