提出MapGlue框架与大规模多模态遥感图像数据集,解决跨模态匹配难题。
MapGlue: Multimodal Remote Sensing Image Matching
- 融合语义上下文与双图引导机制,提取跨模态不变特征
- 在12万+对图像上实现高精度匹配,优于现有方法
- 无需重训练即可泛化到未见模态,适合遥感、导航等应用
多模态遥感图像(MRSI)匹配对跨模态融合、定位和目标检测至关重要,但受成像模态间几何、辐射和视角差异影响,面临严峻挑战。现有单模态数据集规模小、多样性不足,限制了深度学习方案发展。本文提出通用的MRSI匹配框架MapGlue及大规模多模态数据集MapData以填补空白。MapData覆盖全球233个采样点,原始图像尺寸达7,000×5,000至20,000×15,000像素;经严格清洗后,提供121,781对对齐的电子地图-可见光图像(512×512像素),采用混合人工与自动化标注的真值,缓解了可扩展多模态基准稀缺问题。MapGlue通过语义上下文与双图引导机制,实现全局到局部交互,增强描述子对抗模态特异性失真的鲁棒性。在MapData及五个公开数据集上的大量实验表明,该方法在复杂条件下显著优于当前最优模型。尤其值得注意的是,MapGlue无需重新训练即可有效泛化至未见过的模态,展现出强适应能力。本工作通过可扩展数据构建与稳健的语义驱动框架,解决了长期存在的MRSI匹配难题,并在非特定训练任务的其他模态匹配中也表现出优异泛化性能。数据集与代码已开源:https://github.com/PeihaoWu/MapGlue。
原文摘要 · Abstract (English)
Multimodal remote sensing image (MRSI) matching is pivotal for cross-modal fusion, localization, and object detection, but it faces severe challenges due to geometric, radiometric, and viewpoint discrepancies across imaging modalities. Existing unimodal datasets lack scale and diversity, limiting deep learning solutions. This paper proposes MapGlue, a universal MRSI matching framework, and MapData, a large-scale multimodal dataset addressing these gaps. Our contributions are twofold. MapData, a globally diverse dataset spanning 233 sampling points, offers original images (7,000x5,000 to 20,000x15,000 pixels). After rigorous cleaning, it provides 121,781 aligned electronic map-visible image pairs (512x512 pixels) with hybrid manual-automated ground truth, addressing the scarcity of scalable multimodal benchmarks. MapGlue integrates semantic context with a dual graph-guided mechanism to extract cross-modal invariant features. This structure enables global-to-local interaction, enhancing descriptor robustness against modality-specific distortions. Extensive evaluations on MapData and five public datasets demonstrate MapGlue's superiority in matching accuracy under complex conditions, outperforming state-of-the-art methods. Notably, MapGlue generalizes effectively to unseen modalities without retraining, highlighting its adaptability. This work addresses longstanding challenges in MRSI matching by combining scalable dataset construction with a robust, semantics-driven framework. Furthermore, MapGlue shows strong generalization capabilities on other modality matching tasks for which it was not specifically trained. The dataset and code are available at https://github.com/PeihaoWu/MapGlue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。