用掩码和条带卷积提升跨视角物体定位精度
Seeing the Unseen: Mask-Driven Positional Encoding and Strip-Convolution Context Modeling for Cross-View Object Geo-Localization
- 用分割掩码代替关键点,同时捕捉位置和形状信息
- 在卫星图像中对长条形物体定位准确率提升3.39%
- 适合自动驾驶与城市测绘中的高精度定位场景
跨视角物体地理定位可实现高精度物体定位,广泛应用于自动驾驶、城市管理与灾后响应。现有方法依赖基于关键点的位置编码,仅能捕捉二维坐标而忽略物体形状,导致对标注偏移敏感且跨视角匹配能力弱。为此,本文提出基于掩码的位置编码(MPE),利用分割掩码同时捕获空间坐标与物体轮廓,使模型从“位置感知”升级为“物体感知”。针对卫星图像中长条形物体(如延展建筑)的定位难题,设计上下文增强模块(CEM),采用水平与垂直条带卷积核提取长距离上下文特征,增强条状物体间的特征区分能力。融合MPE与CEM,构建端到端框架EDGeo。在两个公开数据集(CVOGL与VIGOR-Building)上实验表明,本方法在复杂地-星视角下定位精度提升3.39%,达到当前最优性能。该工作为跨视角地理定位提供了鲁棒的位置编码范式与上下文建模框架。
原文摘要 · Abstract (English)
Cross-view object geo-localization enables high-precision object localization through cross-view matching, with critical applications in autonomous driving, urban management, and disaster response. However, existing methods rely on keypoint-based positional encoding, which captures only 2D coordinates while neglecting object shape information, resulting in sensitivity to annotation shifts and limited cross-view matching capability. To address these limitations, we propose a mask-based positional encoding scheme that leverages segmentation masks to capture both spatial coordinates and object silhouettes, thereby upgrading the model from "location-aware" to "object-aware." Furthermore, to tackle the challenge of large-span objects (e.g., elongated buildings) in satellite imagery, we design a context enhancement module. This module employs horizontal and vertical strip convolutional kernels to extract long-range contextual features, enhancing feature discrimination among strip-like objects. Integrating MPE and CEM, we present EDGeo, an end-to-end framework for robust cross-view object geo-localization. Extensive experiments on two public datasets (CVOGL and VIGOR-Building) demonstrate that our method achieves state-of-the-art performance, with a 3.39% improvement in localization accuracy under challenging ground-to-satellite scenarios. This work provides a robust positional encoding paradigm and a contextual modeling framework for advancing cross-view geo-localization research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。