通过三维对齐与结构感知,让模型在极端视角下保持语义、结构和几何一致性。
More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

- 构建三维对齐的统一框架,使结构学习成为内在需求。
- 在KITTI和VIGOR数据集上达到最优定位性能。
- 适合关注跨视角理解与空间智能的视觉研究者。
在极端视角变化下实现一致的跨视图理解是空间智能的关键,它使模型能在不同视角间识别同一场景。跨视图定位为这一能力提供了自然路径,要求模型在外观剧烈变化下将地面视图图像与地理参考的卫星视图图像对齐,以估计相机位姿。近期视觉基础模型通过提供丰富的2D表征,使这一长期难题日益可行。然而,我们认为跨视图定位不应仅视为2D匹配或位姿估计。本文重新审视跨视图定位,探究其如何帮助模型在极端视角变化下获得稳定的语义、可靠的结构与可迁移的几何。我们识别出现有方法的三个关键局限:缺乏显式三维支撑、依赖严格的点对点匹配导致语义一致性减弱、学习目标为绝对值而对几何推理引导有限。为此,我们提出CROSS框架,基于三维对齐、结构感知匹配与假设排序,使结构学习成为内在要求,促进语义表示稳定,并实现可迁移几何。在KITTI和VIGOR数据集上的大量实验表明,CROSS在跨视图定位上达到最先进性能,更重要的是,能有效学习跨极端视角的稳定语义、可靠结构与可迁移几何。
原文摘要 · Abstract (English)
Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same scene across extreme viewpoint gaps. Cross-view localization naturally provides a promising pathway toward this ability, as it requires a model to align ground-view imagery with geo-referenced satellite-view imagery despite drastic appearance changes to estimate camera poses. Recent visual foundation models have made this long-standing localization problem increasingly feasible by providing rich 2D representations for cross-view matching. However, we argue that cross-view localization should not be viewed merely as 2D matching or pose estimation. In this work, we revisit cross-view localization as more than pose estimation and investigate how it can help the model develop consistent cross-view understanding under extreme viewpoint changes, including stable semantics, reliable structure, and transferable geometry. We identify three key limitations of existing methods that prevent them from achieving this. They usually lack explicit 3D grounding, rely on strict point-wise matching that can weaken semantic consistency, and learn from an absolute objective that provides limited guidance for geometric reasoning. To address these limitations, we propose CROSS, a unified cross-view localization framework built upon 3D-grounded alignment, structure-aware matching, and hypothesis ranking. This formulation makes structure learning an intrinsic requirement, encourages semantic representations to remain stable, and enables the model to acquire transferable geometry. Extensive experiments on the KITTI and VIGOR datasets show that CROSS achieves state-of-the-art performance in cross-view localization. More importantly, CROSS effectively learns stable semantics, reliable structure, and transferable geometry across extremely different viewpoints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。