无需预设锚点,用位置编码实现跨视角物体精准定位。
Anchor-free Cross-view Object Geo-localization with Gaussian Position Encoding and Cross-view Association
- 直接预测像素到目标框的四向偏移,摆脱锚点依赖。
- 引入高斯位置编码,增强对目标位置的鲁棒性建模。
- 跨视图关联模块提升大外观差异下的定位可靠性。
现有跨视角物体地理定位方法多采用基于锚点的范式,虽有效但受限于预定义锚点。为此,本文提出首个无锚点跨视角物体地理定位方法AFGeo,直接为每个像素预测至真实框的左右上下偏移量,实现无需预设锚点的物体定位。为获取更鲁棒的空间先验,AFGeo引入高斯位置编码(GPE)建模查询图像中的点击点,缓解跨视角场景下物体位置不确定带来的挑战。此外,模型设计了跨视图物体关联模块(CVOAM),在不同视角间建立同一物体及其上下文的关联,提升在显著外观差异下的定位可靠性。通过结合无锚点定位、GPE与CVOAM,AFGeo以极低参数开销实现轻量高效,且在基准数据集上达到领先性能。
原文摘要 · Abstract (English)
Most existing cross-view object geo-localization approaches adopt anchor-based paradigm. Although effective, such methods are inherently constrained by predefined anchors. To eliminate this dependency, we first propose an anchor-free formulation for cross-view object geo-localization, termed AFGeo. AFGeo directly predicts the four directional offsets (left, right, top, bottom) to the ground-truth box for each pixel, thereby localizing the object without any predefined anchors. To obtain a more robust spatial prior, AFGeo incorporates Gaussian Position Encoding (GPE) to model the click point in the query image, mitigating the uncertainty of object position that challenges object localization in cross-view scenarios. In addition, AFGeo incorporates a Cross-view Object Association Module (CVOAM) that relates the same object and its surrounding context across viewpoints, enabling reliable localization under large cross-view appearance gaps. By adopting an anchor-free localization paradigm that integrates GPE and CVOAM with minimal parameter overhead, our model is both lightweight and computationally efficient, achieving state-of-the-art performance on benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。