让视觉语言模型学会自主进化地理推理能力。
Skill-Conditioned Visual Geolocation for Vision-Language Models

- 构建可生长的技能图谱,用自然语言技能引导地理定位推理。
- 在GeoRC数据集上准确率提升,且推理更可信、泛化更强。
- 无需参数更新,自动从网络数据中学习新技能并修正偏差。
视觉语言模型(VLMs)在图像地理定位方面展现出潜力,但仍缺乏结构化地理推理能力和自主进化能力。现有方法多依赖隐式参数记忆,常使用过时知识并产生幻觉推理。当前推理为一次性过程,缺乏基于推理结果的反馈循环。为此,我们提出GeoSkill——一种无需训练的框架,基于可演化的技能图谱(Skill-Graph)。首先,将人类专家轨迹提炼为原子级自然语言技能;推理时,由推理模型直接依据当前技能图谱进行导航;为持续成长,引入自主演化机制,利用大模型对来自网络规模数据的图像-坐标对进行多次推理回溯,并通过真实世界验证结果,分析成功与失败轨迹,迭代合成与剪枝技能,从而在不更新参数的前提下扩展技能图谱并纠正地理偏差。实验表明,GeoSkill在GeoRC数据集上实现优异的定位精度与推理忠实性,并在多种外部数据集上保持强泛化能力。此外,自主演化促使新型可验证技能涌现,显著增强系统对真实地理知识的认知,超越孤立案例研究。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have shown a promising ability in image geolocation, but they still lack structured geographic reasoning and the capacity for autonomous self-evolution. Existing methods predominantly rely on implicit parametric memory, which often exploits outdated knowledge and generates hallucinated reasoning. Furthermore, current inference is a "one-off" process, lacking the feedback loops necessary for self-evolution based on reasoning outcomes. To address these issues, we propose GeoSkill, a training-free framework based on an evolving Skill-Graph. We first initialize the graph by refining human expert trajectories into atomic, natural-language skills. For execution, GeoSkill employs an inference model to perform direct reasoning guided by the current Skill-Graph. For continuous growth, an Autonomous Evolution mechanism leverages a larger model to conduct multiple reasoning rollouts on image-coordinate pairs sourced from web-scale data and verified real-world reasoning. By analyzing both successful and failed trajectories from these rollouts, the mechanism iteratively synthesizes and prunes skills, effectively expanding the Skill-Graph and correcting geographic biases without any parameter updates. Experiments demonstrate that GeoSkill achieves promising performance in both geolocation accuracy and reasoning faithfulness on GeoRC, while maintaining superior generalization across diverse external datasets. Furthermore, our autonomous evolution fosters the emergence of novel, verifiable skills, significantly enhancing the system's cognition of real-world geographic knowledge beyond isolated case studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。