用视点距离掩码提升卫星图像匹配精度
EpiMask: Leveraging Epipolar Distance Based Masks in Cross-Attention for Satellite Image Matching
- 基于视点距离设计注意力掩码,限制匹配区域
- 在SatDepth数据集上精度提升30%
- 适合需要高精度卫星图像匹配的研究者
基于深度学习的图像匹配网络能处理较大视角和光照变化,实现亚像素级像素对齐。这些网络通常在地面图像数据集上训练,隐式优化于针孔相机几何结构,因此在卫星图像匹配中表现不佳——因为卫星图像是移动卫星逐行扫描生成的。本文提出EpiMask,一种用于卫星图像的半密集匹配网络:(1) 使用局部仿射近似建模相机几何;(2) 采用基于视点距离的注意力掩码,仅在几何上合理的区域进行交叉注意力;(3) 对预训练图像编码器进行微调以增强特征鲁棒性。在SatDepth数据集上的实验表明,相比重新训练的地面模型,匹配精度最高提升30%。
原文摘要 · Abstract (English)
The deep-learning based image matching networks can now handle significantly larger variations in viewpoints and illuminations while providing matched pairs of pixels with sub-pixel precision. These networks have been trained with ground-based image datasets and, implicitly, their performance is optimized for the pinhole camera geometry. Consequently, you get suboptimal performance when such networks are used to match satellite images since those images are synthesized as a moving satellite camera records one line at a time of the points on the ground. In this paper, we present EpiMask, a semi-dense image matching network for satellite images that (1) Incorporates patch-wise affine approximations to the camera modeling geometry; (2) Uses an epipolar distance-based attention mask to restrict cross-attention to geometrically plausible regions; and (3) That fine-tunes a foundational pretrained image encoder for robust feature extraction. Experiments on the SatDepth dataset demonstrate up to 30% improvement in matching accuracy compared to re-trained ground-based models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。