arXiv:2608.15251cs.CV2026-08

无需检测器的特征匹配+多视角轨迹优化,提升航地图像三维重建精度

Robust structure from motion for aerial-ground images via detector-free feature matching and multi-view track refinement

论文配图:Robust structure from motion for aerial-ground images via detector-free feature matching and multi-view track refinement
图 1 · 摘自论文原文
  • 用旋转感知模块替代卷积,生成不变于旋转的特征图
  • 在真实数据集上比LoFTR提升93.9%的5°姿态误差AUC
  • 适合需要高精度航地融合三维建模的场景

航地图像融合三维重建对生成高精度城市模型至关重要,但视角、尺度和旋转差异导致特征匹配极难。本文提出一种旋转鲁棒的无检测器匹配网络与多视角轨迹优化方法,用于增量式结构光(ISfM)。流程包含四个模块:首先,采用全向状态空间块(OSS Block)替代传统卷积,通过八个对称方向选择性扫描,建模长程空间依赖并生成旋转不变特征图;其次,利用四叉树注意力构建分层令牌金字塔,以线性复杂度捕捉长程上下文,隔离高关联区域并剔除无关部分;第三,双向特征匹配采用对称粗到精策略,粗匹配在互近邻约束下计算双方向Softmax置信矩阵,细匹配使用多层感知机回归亚像素偏移;最后,多视角轨迹优化通过集成索引结构评估局部空间邻近性,将断裂子轨迹链接至最高置信锚点,保障整条流水线中特征重复性稳定。基于真实航地数据集实验表明,该方法在5°姿态误差下的AUC相比LoFTR提升93.9%,在ISfM重建中达到最高精度,准确率提高27.6%至32.7%。所提方法为航地图像融合三维重建提供了可靠解决方案。

原文摘要 · Abstract (English)

Integrated 3D reconstruction from aerial-ground images is essential for generating high-precision urban 3D models, yet severe variations in viewpoint, scale, and rotation make robust feature matching highly challenging. To address these limitations, this study introduces a rotation-robust detector-free matching network coupled with multi-view track refinement for incremental Structure from Motion (ISfM). The proposed workflow features four key modules. First, rotation-aware feature extraction replaces traditional convolutions with an Omnidirectional State Space Block (OSS Block) that selectively scans across eight symmetrical directions to model long-range spatial dependencies and synthesize rotation-invariant feature maps. Second, multi-scale attention transformation utilizes quadtree attention to build a hierarchical token pyramid that isolates high-association token regions and discards irrelevant areas, capturing long-range context with linear computational complexity. Third, bi-directional feature matching executes a symmetric coarse-to-fine matching scheme where coarse alignment computes dual-direction Softmax confidence matrices under mutual nearest neighbor constraints, and fine alignment uses a multi-layer perceptron to regress sub-pixel coordinate offsets. Finally, multi-view track refinement employs an integrated indexing structure to evaluate localized spatial proximity and link disjoint sub-tracks to the highest-confidence anchor point, ensuring stable feature repeatability across the ISfM pipeline. By using real aerial-ground datasets, experimental results demonstrate that the proposed method improves AUC at 5° pose error by 93.9% compared with LoFTR and achieves the highest precision in ISfM reconstruction, with the improved accuracy ranging from 27.6% to 32.7%. The proposed method provides a reliable solution for integrated 3D reconstruction of aerial-ground images.

三维重建航地融合特征匹配结构光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。