arXiv:2608.04106cs.CVeess.IV2026-08

LoRetta解决全球遥感图像密集匹配难题,提升精度与效率。

LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching

论文配图:LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching
图 1 · 摘自论文原文
  • 先定位可匹配区域和几何关系,再精细对齐,提升匹配可靠性。
  • 在LEVIR-GM数据集上AUC达83.3%,比最强基线提升1.6点,延迟降低47.8%。
  • 适用于跨时间、跨视角遥感图像配准,适合地理信息与测绘研究者。

密集图像匹配建立像素级对应关系,在计算机视觉与摄影测量中应用广泛。然而,将密集匹配扩展至全球尺度的遥感影像仍具挑战,因影像对可能在获取时间、季节、视角、空间分辨率及地表状态上存在差异。由此导致的大范围几何偏移、部分重叠及固有不可匹配区域使直接预测密集对应关系不可靠且低效。为此,本文将密集匹配重新定义为定位与注册:先定位可匹配重叠区域及仿射几何关系,再在对齐区域内精修密集残差。基于此,提出LoRetta基础模型,结合可匹配性感知的仿射定位与引导式密集注册。同时引入LEVIR-GM,一个覆盖六大陆、五年跨度、0.5–1024米分辨率的全球多时相光学匹配基准数据集,包含10.3万对原始对齐样本与82.7万对增强样本,并提供原生可匹配性标签。进一步建立了统一评估协议,涵盖稀疏、半密集与密集匹配器。在LEVIR-GM上,LoRetta AUC达83.3%,优于最强基线RoMa v2的1.6个百分点,1像素与2像素正确关键点率(PCK)分别提升6.5与8.2点,推理延迟降低47.8%。宇航员到卫星及无人机到卫星地理定位实验进一步验证其作为可复用几何对齐器的迁移能力。

原文摘要 · Abstract (English)

Dense image matching establishes pixel-wise correspondences and underpins broad applications in computer vision and photogrammetry. However, extending dense matching to global-scale remote sensing remains challenging because image pairs may differ in acquisition time, season, viewpoint, spatial resolution, and land-cover state. The resulting large geometric offsets, partial overlap, and intrinsically unmatchable regions make direct dense correspondence prediction unreliable and inefficient. We thus reformulate dense matching as localization-and-registration: first localizing the matchable overlap and affine geometry, then refining dense residuals within the aligned frame. Based on this formulation, we propose LoRetta, a foundation model coupling matchability-aware affine localization with guided dense registration. We also introduce LEVIR-GM, a global-scale multi-temporal optical matching benchmark with dataset-native matchability labels (103K aligned, 827K augmented pairs, six continents, five years, 0.5-1024 m resolution). We further establish a unified evaluation protocol for sparse, semi-dense, and dense matchers. On LEVIR-GM, LoRetta achieves an area under the curve (AUC) of 83.3%, outperforming the strongest baseline RoMa v2 by 1.6 points, with larger percentage of correct keypoints (PCK) gains of 6.5 and 8.2 points at 1 and 2 pixels, while reducing inference latency by 47.8%. Astronaut-to-satellite and unmanned aerial vehicle (UAV)-to-satellite geolocalization experiments further demonstrate its transferability as a reusable geometric aligner.

遥感图像匹配基础模型地理定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。