通过融合时空特征提升遥感图像跨模态匹配精度,显著改善目标检测效果。
Cross Spatial Temporal Fusion Attention for Remote Sensing Object Detection via Image Feature Matching
- 引入跨时空融合机制,利用关键点匹配构建多区域对应关系
- 在HRSC2016和DOTA上分别达到90.99%和90.86%的mAP,性能领先
- 兼顾局部细节与上下文信息,适合遥感图像目标检测场景
由于多模态遥感图像间存在显著的几何与辐射差异,有效描述跨模态图像匹配特征仍具挑战。现有方法多在全连接层提取特征,难以充分捕捉跨模态相似性。本文提出跨时空融合(CSTF)机制,通过独立检测参考图与查询图中的尺度不变关键点,增强特征表示。该方法通过两种方式提升匹配性能:一是构建利用多个图像区域信息的对应关系图;二是将相似性匹配重构为使用SoftMax与全卷积网络(FCN)层的分类任务。双重设计使CSTF既能敏感捕捉独特局部特征,又融入全局上下文信息,实现对多种遥感模态的鲁棒匹配。为验证改进匹配的实际效用,我们在HRSC2016与DOTA基准数据集上评估了其在目标检测任务中的表现。结果表明,该方法在HRSC2016上取得90.99%的平均mAP,DOTA上达90.86%,超越现有模型。同时,模型推理速度保持在12.5 FPS,计算高效。实证显示,该跨模态特征匹配方法可直接提升下游遥感应用如目标检测的性能。
原文摘要 · Abstract (English)
Effectively describing features for cross-modal remote sensing image matching remains a challenging task due to the significant geometric and radiometric differences between multimodal images. Existing methods primarily extract features at the fully connected layer but often fail to capture cross-modal similarities effectively. We propose a Cross Spatial Temporal Fusion (CSTF) mechanism that enhances feature representation by integrating scale-invariant keypoints detected independently in both reference and query images. Our approach improves feature matching in two ways: First, by creating correspondence maps that leverage information from multiple image regions simultaneously, and second, by reformulating the similarity matching process as a classification task using SoftMax and Fully Convolutional Network (FCN) layers. This dual approach enables CSTF to maintain sensitivity to distinctive local features while incorporating broader contextual information, resulting in robust matching across diverse remote sensing modalities. To demonstrate the practical utility of improved feature matching, we evaluate CSTF on object detection tasks using the HRSC2016 and DOTA benchmark datasets. Our method achieves state-of-theart performance with an average mAP of 90.99% on HRSC2016 and 90.86% on DOTA, outperforming existing models. The CSTF model maintains computational efficiency with an inference speed of 12.5 FPS. These results validate that our approach to crossmodal feature matching directly enhances downstream remote sensing applications such as object detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。