arXiv:2511.02489cs.CV2025-11被引 4

针对红外可见光遥感图像定位难题,提出基于目标感知的图匹配方法。

Object-aware graph matching network for cross-domain remote sensing image localization

  • 融合目标检测与双图神经网络,建模图像内外部结构关系
  • 在IRVL328数据集上实现强于现有方法的跨模态定位性能
  • 适合研究跨域遥感图像配准、多源传感器融合的学者参考

跨域与跨模态遥感图像地理定位仍面临巨大挑战,源于异构传感器与平台间显著的外观差异和不稳定的语义对应关系。现有方法主要依赖场景级或全局表征,在复杂环境下难以实现可靠对齐,尤其在红外到可见光匹配等严重模态差异情况下表现不佳。为此,本研究构建了新的红外-可见光遥感定位数据集IRVL328,以反映真实场景中的挑战性跨模态变化。同时提出一种目标感知的图匹配框架,将显著结构区域表示为图节点,联合建模跨图像对应关系与图像内关系;引入仅训练阶段使用的节点对齐策略,在不增加推理复杂度的前提下增强监督。实验结果表明,该方法在SUES-200上表现优异,在IRVL328和DenseUAV上尤其突出,尤其在跨模态设置中展现强大鲁棒性。

原文摘要 · Abstract (English)

Cross-domain and cross-modal remote sensing image geo-localization remains challenging due to large appearance discrepancies and unstable semantic correspondence across heterogeneous sensors and platforms. Existing methods mainly rely on scene-level or global representations, which often struggle to achieve reliable alignment in complex environments, especially under severe modality gaps such as infrared-to-visible matching. To facilitate research in this setting, this study introduces IRVL328, a new infrared-visible remote sensing localization dataset designed to reflect challenging cross-modal variations. Meanwhile, this study proposes an object-aware graph matching framework that integrates object detection with dual-graph neural reasoning, where salient structural regions are represented as graph nodes and both inter-image correspondences and intra-image relations are jointly modeled; a training-only node alignment strategy is further introduced to enhance supervision without increasing inference complexity. Experimental results show that the proposed method achieves competitive performance on SUES-200 and strong performance on IRVL328 and DenseUAV, particularly in challenging cross-modal settings.

遥感图像跨模态图神经网络目标感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。