用扩散模型直接预测多视角图像的三维结构和相机位姿
DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion
- 将相机位姿与场景结构建模为全局坐标系中的射线起点和终点
- 在合成与真实数据集上均优于传统及学习型方法,且能自然建模不确定性
- 专设机制应对缺失数据和无界坐标问题,适合需要鲁棒重建的场景
当前结构从运动(SfM)方法通常采用两阶段流程,结合学习或几何的成对推理与后续全局优化。本文提出一种数据驱动的多视图推理方法,可直接从多视图图像中推断三维场景几何与相机位姿。DiffusionSfM 将场景几何与相机参数化为全局坐标系下的像素级射线起点与终点,并利用基于 Transformer 的去噪扩散模型从多视图输入中预测它们。为应对训练中缺失数据与无界场景坐标的实际挑战,我们引入专用机制以保障稳健学习。我们在合成与真实数据集上进行实证验证,结果表明 DiffusionSfM 在性能上超越经典与学习型方法,且能自然建模不确定性。
原文摘要 · Abstract (English)
Current Structure-from-Motion (SfM) methods typically follow a two-stage pipeline, combining learned or geometric pairwise reasoning with a subsequent global optimization step. In contrast, we propose a data-driven multi-view reasoning approach that directly infers 3D scene geometry and camera poses from multi-view images. Our framework, DiffusionSfM, parameterizes scene geometry and cameras as pixel-wise ray origins and endpoints in a global frame and employs a transformer-based denoising diffusion model to predict them from multi-view inputs. To address practical challenges in training diffusion models with missing data and unbounded scene coordinates, we introduce specialized mechanisms that ensure robust learning. We empirically validate DiffusionSfM on both synthetic and real datasets, demonstrating that it outperforms classical and learning-based approaches while naturally modeling uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。