arXiv:2604.10217cs.CV2026-04被引 2

评估预训练图像匹配器在遥感雷达-光学图像配准中的表现,发现无需微调也能达到高精度。

Are Pretrained Image Matchers Good Enough for SAR-Optical Satellite Registration?

  • 零样本测试24种预训练匹配器,用大图分块推理与几何过滤提升鲁棒性
  • RoMa和XoFTR分别实现3.0像素和3.4像素的最低平均同名点误差
  • 模型部署时几何模型和分块策略选择影响可达33倍误差变化,远超模型本身差异

跨模态光学—合成孔径雷达(SAR)配准是遥感灾害响应中的瓶颈。现代图像匹配器几乎仅在自然图像域上开发和评测。本文在零样本设置下评估了24种预训练匹配器配置,未对卫星或SAR数据进行微调或领域自适应。评估涵盖SpaceNet9及另外两个跨模态基准,采用确定性协议:使用分块大图推理、鲁棒几何滤波及基于同名点的度量标准。结果表明迁移性能不均:具有显式跨模态训练的匹配器并未普遍优于无此训练者。XoFTR(为可见光—热成像匹配训练)和RoMa在标注的SpaceNet9训练场景中取得最低报告均值同名点误差,分别为3.0像素和3.4像素。其中,RoMa未经过任何跨模态训练。此外,XoFTR每对图像推理时间仅0.4秒,比RoMa家族快约一个数量级(5.0秒)。RoMa的表现支持冻结DINOv2特征可增强对显著外观变化的鲁棒性假设。部署协议选择(几何模型、分块大小、内点门控)使单个匹配器的均值误差最高变动达33倍,该变化甚至超过在同一协议下更换匹配器的影响。仅使用仿射几何模型即可将均值误差从12.3像素降至9.7像素。

原文摘要 · Abstract (English)

Cross-modal optical--SAR (Synthetic Aperture Radar) registration is a bottleneck in remote-sensing disaster response. Modern image matchers are developed and benchmarked almost exclusively on natural-image domains. We evaluate twenty-four pretrained matcher configurations in a zero-shot setting, with no fine-tuning or domain adaptation on satellite or SAR data. The evaluation spans SpaceNet9 and two additional cross-modal benchmarks under a deterministic protocol that uses tiled large-image inference, robust geometric filtering, and tie-point-grounded metrics. Our results show uneven transfer: matchers with explicit cross-modal training do not uniformly outperform those without it. XoFTR (trained for visible--thermal matching) and RoMa achieve the lowest reported mean tie-point error at $3.0$ px on the labeled SpaceNet9 training scenes. RoMa achieves this result \emph{without any cross-modal training}. MatchAnything-ELoFTR ($3.4$ px), trained on synthetic cross-modal pairs, is close behind. XoFTR also runs roughly an order of magnitude faster than the RoMa family ($0.4$ s vs. $5.0$ s per pair). RoMa's performance is consistent with the hypothesis that frozen DINOv2 features confer robustness to large appearance shifts. Deployment protocol choices (geometry model, tile size, inlier gating) change mean error by up to $33\times$ for a single matcher. This shift can exceed the effect of swapping matchers within our evaluated protocol ablations. Affine geometry alone reduces mean error from $12.3$ to $9.7$ px.

图像匹配遥感SAR零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。