提出双注意力机制提升跨视角物体定位精度
Improving Cross-view Object Geo-localization: A Dual Attention Approach with Cross-view Interaction and Multi-Scale Spatial Features
- 设计跨视图交互模块实现双向信息迭代
- 多尺度空间注意力增强特征表达,抑制无关噪声
- 构建G2D数据集填补地面到无人机定位空白
跨视角物体地理定位因潜在应用价值受到关注。现有方法通过注意力机制捕捉不同视角间查询物体的空间依赖关系,生成空间关系特征图以预测位置。然而,这些方法未能有效实现视图间信息传递,且未对空间关系特征图进行进一步优化,导致模型错误聚焦于无关边缘噪声,影响定位性能。为此,本文提出跨视图交叉注意力模块(CVCAM),通过多次迭代的视图间交互,实现查询物体上下文信息的持续交换与学习,深化对跨视图关系的理解,同时抑制与查询物体无关的边缘噪声。此外,引入多头空间注意力模块(MHSAM),利用不同尺寸卷积核从蕴含隐式对应关系的特征图中提取多尺度空间特征,进一步增强查询物体的特征表示。针对跨视角物体地理定位数据集稀缺问题,构建了名为G2D的新数据集,用于“地面到无人机”定位任务,丰富了现有数据资源。在CVOGL和G2D数据集上的大量实验表明,所提方法在定位精度上显著优于当前最优水平。
原文摘要 · Abstract (English)
Cross-view object geo-localization has recently gained attention due to potential applications. Existing methods aim to capture spatial dependencies of query objects between different views through attention mechanisms to obtain spatial relationship feature maps, which are then used to predict object locations. Although promising, these approaches fail to effectively transfer information between views and do not further refine the spatial relationship feature maps. This results in the model erroneously focusing on irrelevant edge noise, thereby affecting localization performance. To address these limitations, we introduce a Cross-view and Cross-attention Module (CVCAM), which performs multiple iterations of interaction between the two views, enabling continuous exchange and learning of contextual information about the query object from both perspectives. This facilitates a deeper understanding of cross-view relationships while suppressing the edge noise unrelated to the query object. Furthermore, we integrate a Multi-head Spatial Attention Module (MHSAM), which employs convolutional kernels of various sizes to extract multi-scale spatial features from the feature maps containing implicit correspondences, further enhancing the feature representation of the query object. Additionally, given the scarcity of datasets for cross-view object geo-localization, we created a new dataset called G2D for the "Ground-to-Drone" localization task, enriching existing datasets and filling the gap in "Ground-to-Drone" localization task. Extensive experiments on the CVOGL and G2D datasets demonstrate that our proposed method achieves high localization accuracy, surpassing the current state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。