arXiv:2508.00152cs.CV2025-08ICCV被引 3

用好奇心驱动探索提升地理定位的鲁棒性

GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration

  • 引入内在奖励机制,让智能体主动探索未知区域
  • 在四个基准上均优于传统方法,尤其对陌生目标表现更好
  • 适合需要强泛化能力的地理定位场景

主动地理定位(AGL)旨在将目标(如航拍图、地面图像或文本)在预定义搜索区域内精确定位。现有方法将AGL视为基于距离奖励的强化学习问题,通过隐式学习最小化与目标的相对距离来定位。然而,当距离估计困难或遇到未见过的目标与环境时,由于训练中探索策略不可靠,智能体的鲁棒性和泛化能力下降。本文提出GeoExplorer,一种通过内在奖励实现好奇心驱动探索的AGL智能体。其奖励机制不依赖具体目标,可基于有效环境建模实现稳健、多样且上下文相关的探索。在四个AGL基准上的大量实验表明,GeoExplorer在多样化场景中具有优异的有效性与泛化能力,尤其擅长定位未见过的目标与环境。

原文摘要 · Abstract (English)

Active Geo-localization (AGL) is the task of localizing a goal, represented in various modalities (e.g., aerial images, ground-level images, or text), within a predefined search area. Current methods approach AGL as a goal-reaching reinforcement learning (RL) problem with a distance-based reward. They localize the goal by implicitly learning to minimize the relative distance from it. However, when distance estimation becomes challenging or when encountering unseen targets and environments, the agent exhibits reduced robustness and generalization ability due to the less reliable exploration strategy learned during training. In this paper, we propose GeoExplorer, an AGL agent that incorporates curiosity-driven exploration through intrinsic rewards. Unlike distance-based rewards, our curiosity-driven reward is goal-agnostic, enabling robust, diverse, and contextually relevant exploration based on effective environment modeling. These capabilities have been proven through extensive experiments across four AGL benchmarks, demonstrating the effectiveness and generalization ability of GeoExplorer in diverse settings, particularly in localizing unfamiliar targets and environments.

地理定位强化学习好奇心驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。