提出统一单阶段框架,实现跨视角物体精准定位与空间建模。
Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization

- 基于3D先验的单阶段模型,融合视觉特征与多模态提示。
- 在22万对图像上达到领先性能,零样本跨视角定位效果好。
- 适合需要高精度地理定位与多视图泛化的研究者使用。
跨视角物体地理定位旨在从查询视图(如地面或无人机)中定位目标物体,对应到带有地理标签的参考图像(如卫星图)。现有方法严重依赖2D外观匹配,受限于缺乏几何元数据、多样提示和标准视场图像的小规模数据集。为此,我们首次引入 extit{dataset},一个大规模、高保真建筑数据集,包含超过22万对地面-卫星及无人机-卫星图像对,提供多模态提示(点、框、掩码)和相机位姿,支持灵活的目标指代与显式空间建模。此外,我们提出一种新型单阶段几何感知定位框架GAGeo,基于排列等变3D基础模型$π^3$,通过无缝整合视觉特征、指代提示与可学习任务标记,在一次前向传播中联合预测边界框、分割掩码和相机位姿。同时,我们设计了一种对比损失,以卫星视图为通用锚点,隐式对齐地面与无人机表征,实现无需三元组训练数据的零样本地面到无人机定位。大量实验表明,该方法显著优于现有最优方法,在未见场景和新颖跨视角设置下展现出卓越泛化能力。
原文摘要 · Abstract (English)
Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.g., ground or drone) within a geo-tagged reference image (e.g., satellite). Existing approaches heavily rely on 2D appearance matching and are constrained by limited datasets lacking geometric metadata, diverse prompts, and standard field-of-view imagery. To address these intertwined challenges, we first introduce \dataset, a large-scale, high-fidelity building dataset comprising over 220,000 ground-satellite and drone-satellite pairs. It provides multi-modal prompts (points, boxes, masks) and camera poses to enable flexible target referring and explicit spatial modeling. Furthermore, we propose a novel single-stage Geometry-Aware Geo-localization framework (GAGeo), built upon the permutation-equivariant 3D foundation model $π^3$. By seamlessly integrating visual features, referring prompts, and learnable task tokens, our model adapts the inherited 3D prior to jointly predict bounding boxes, segmentation masks, and camera poses in a single forward pass. Additionally, we introduce a contrastive loss that utilizes the satellite view as a universal anchor, implicitly aligning ground and drone representations to enable zero-shot ground-to-drone localization without requiring triplet training data. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods, exhibiting exceptional generalization ability in unseen scenes and novel cross-view setups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。