利用平面几何关系提升相机位姿估计,解决传统方法在平场景下的失效问题。
Planar-SfM: Camera Pose Estimation via Homography Graph Embeddings

- 通过多视角平面间的单应性关系,构建相机位姿的独立估计
- 在篮球场等高度平面场景中表现显著优于现有方法,户外场景也达到顶尖水平
- 基于图嵌入与谱分析,自动筛选可靠位姿连接,适合复杂平面环境
结构光从运动(SfM)系统在平场景中通常表现不佳,因标准对极几何方法在此类场景中退化。本文提出一种统一框架,将平面表面视为几何约束的来源而非限制。关键洞察在于:任意可见于多个视图的平面可提供相对相机位姿的独立估计,通过同源性分解实现。通过聚合多个平面或单一主导平面的估计,可在传统方法失效的场景中实现鲁棒位姿恢复。我们引入一种新型图基方法,从同源性估计构建位姿图,并采用谱嵌入识别与过滤不可靠边。该方法根据几何与视觉一致性,将同源性位姿估计映射至实线,高效提取最大一致支撑树以完成位姿恢复。本方法天然适用于高度平面场景(如室内体育场馆)及一般三维环境。实验表明,在篮球场图像上性能显著超越现有方法,同时在IMC Phototourism基准的非受限户外场景中达到或超过当前最优水平。
原文摘要 · Abstract (English)
Structure from Motion (SfM) systems traditionally struggle with planar scenes, where standard epipolar geometry-based methods become degenerate. Rather than viewing planar surfaces as a limitation, we propose a unified framework that leverages them as a source of geometric constraints. Our key insight is that each planar surface visible across multiple views provides an independent estimate of relative camera poses through homography decomposition. By aggregating estimates from multiple planes or even from a single dominant plane we achieve robust pose recovery in scenarios where traditional methods fail. We introduce a novel graph-based approach that constructs a pose-graph from homography estimates and employs spectral embedding to identify and filter unreliable edges. Our method maps homography-based pose estimates onto the real line based on their geometric and visual consistency, enabling efficient extraction of a maximally consistent spanning tree for pose recovery. This approach naturally handles both highly planar scenes, such as indoor sports arenas, and general $3$D environments. We demonstrate superior performance on basketball court imagery where existing methods struggle, while matching or exceeding state-of-the-art results on unconstrained outdoor scenes from the IMC Phototourism benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。