无需训练和GPU,精准对齐大尺度卫星影像。
Self-Calibrating Dense Displacement Fields for Reliable Co-Registration of Large Optical Satellite Imagery

- 用像素级位移场自校准,覆盖复杂场景运动
- 在真实数据上将误差中位数降至4.17米
- 单核CPU处理8192²图像,适合无资源环境
多时相与多传感器光学卫星影像的配准依赖于精确对齐,但现有产品仍存在远超像素级的偏移,导致变化检测、时间序列分析和数据融合性能下降。真实图像对在传感器响应、场景内容、观测几何、分辨率及拼接缝等多个维度同时存在差异,且拼接缝处非全局运动。现有工具依赖预设运动模型与调参常数,不匹配时失败或返回错误结果却无提示;学习型匹配器需GPU且泛化能力有限。本文提出SCDF(自校准密集位移场),无需训练与GPU,以像素级位移场为自身运动模型,确保无遗漏。通过金字塔结构的预测-测量-过滤循环:累计位移场预测移动图像块在参考图中的位置,RootSIFT匹配与相关性计算实现亚像素精度测量,所有滤波阈值均基于图像对自校准。单一配置无需数据集调优,可在单个CPU核心上处理8192²图像。在由真实Sentinel-2、Landsat-8/9和NAIP影像构建的584组带真值的数据对上,优于七种经典基线与两种零样本预训练匹配器,实现零失败配准,将最优基线的中位端点误差从6.83米降至4.17米,90%分位误差从17.8米降至7.77米。
原文摘要 · Abstract (English)
Co-registration underlies nearly every multi-temporal and multi-sensor use of optical satellite imagery, and operational products still carry documented offsets well above the fraction-of-a-pixel scale at which change detection, time series, and data fusion degrade. Real image pairs differ along several axes at once (sensor response, scene content, viewing geometry, resolution, mosaic seams), and the last of these is not a single global motion. Existing tools embed a motion model and constants tuned to their development data; a pair that fits is registered precisely, while one that does not either fails to match or returns a result wrong by tens of pixels with no failure reported. Learned matchers add a GPU requirement and carry no accuracy guarantee outside their training distribution. We present SCDF (self-calibrating displacement fields), a training-free, GPU-free estimator whose motion model is the dense per-pixel displacement field itself, so no scene motion falls outside the model. A single predict--measure--filter loop runs over a resolution pyramid: the accumulated field predicts where each patch of the moving image falls in the reference, RootSIFT matching and a correlation pass measure the displacement there to sub-pixel precision, and filters whose thresholds are all calibrated on the image pair itself decide what survives. One configuration, with no per-dataset tuning, processes full $8192^2$ scenes on a single CPU core. On 584 constructed-ground-truth pairs built from real Sentinel-2, Landsat-8/9, and NAIP imagery, against seven classical baselines and two zero-shot pretrained matchers, SCDF registers every pair with zero failures, reduces the best baseline's real-pair median end-point error from 6.83 to 4.17m, and cuts its 90th percentile from 17.8 to 7.77m.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。