用视觉推理实现无GPS的地面到航拍定位
Lifting Vision: Ground to Aerial Localization with Reasoning Guided Planning
- 基于视觉表示进行逐步推理与路径规划
- 在跨视角检索中提升空间推理准确率
- 适合无人导航、隐私保护等场景
多模态智能在视觉理解与高层推理方面取得显著进展,但多数推理系统仍依赖文本信息进行推断,限制了其在视觉导航和地理定位等空间任务中的表现。本文探讨该领域潜力,提出一种新型视觉推理范式——地理一致视觉规划(Geo-Consistent Visual Planning),并构建名为ViReLoc的框架,仅使用视觉表示完成规划与定位。该框架学习空间依赖关系与几何结构,通过强化学习目标优化,实现从地面图像到航拍图像的路径规划。系统融合对比学习与自适应特征交互,对齐多视角观点,减少视角差异。在多样化导航与定位场景中,实验显示其在空间推理准确率与跨视图检索性能上均持续提升。结果表明,视觉推理可作为导航与定位的有效补充方法,且无需实时全球定位系统数据,为更安全的导航方案提供可能。
原文摘要 · Abstract (English)
Multimodal intelligence development recently show strong progress in visual understanding and high level reasoning. Though, most reasoning system still reply on textual information as the main medium for inference. This limit their effectiveness in spatial tasks such as visual navigation and geo-localization. This work discuss about the potential scope of this field and eventually propose an idea visual reasoning paradigm Geo-Consistent Visual Planning, our introduced framework called Visual Reasoning for Localization, or ViReLoc, which performs planning and localization using only visual representations. The proposed framework learns spatial dependencies and geometric relations that text based reasoning often suffer to understand. By encoding step by step inference in the visual domain and optimizing with reinforcement based objectives, ViReLoc plans routes between two given ground images. The system also integrates contrastive learning and adaptive feature interaction to align cross view perspectives and reduce viewpoint differences. Experiments across diverse navigation and localization scenarios show consistent improvements in spatial reasoning accuracy and cross view retrieval performance. These results establish visual reasoning as a strong complementary approach for navigation and localization, and show that such tasks can be performed without real time global positioning system data, leading to more secure navigation solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。