通过多智能体结构化辩论提升图像地理定位精度
GraphGeo: Multi-Agent Debate Framework for Visual Geo-localization with Heterogeneous Graph Neural Networks
- 用异构图神经网络建模智能体间的协作、竞争与知识传递关系
- 在多个基准上超越现有最优方法,显著提升定位准确率
- 适合需要高精度地理定位的导航与遥感应用
视觉地理定位需依赖广泛地理知识和复杂推理来确定无GPS元数据图像的位置。传统检索方法受限于数据库覆盖范围与质量。近期大型视觉语言模型(LVLMs)可直接从图像内容推断位置,但单一模型在多样地理区域和复杂场景下表现不佳。现有多智能体系统虽通过模型协同提升性能,却对所有交互一视同仁,缺乏有效处理预测冲突的机制。我们提出GraphGeo,一种基于异构图神经网络的多智能体辩论框架,用于视觉地理定位。该方法通过类型化边建模多样化辩论关系,区分支持性协作、竞争性论辩与知识传递。引入节点级精炼与边级论辩建模的双层辩论机制,并设计跨层级拓扑优化策略,实现图结构与智能体表征的共同演化。在多个基准上的实验表明,GraphGeo显著优于当前最先进方法。该框架将智能体间的认知冲突转化为结构化辩论,从而提升地理定位准确性。
原文摘要 · Abstract (English)
Visual geo-localization requires extensive geographic knowledge and sophisticated reasoning to determine image locations without GPS metadata. Traditional retrieval methods are constrained by database coverage and quality. Recent Large Vision-Language Models (LVLMs) enable direct location reasoning from image content, yet individual models struggle with diverse geographic regions and complex scenes. Existing multi-agent systems improve performance through model collaboration but treat all agent interactions uniformly. They lack mechanisms to handle conflicting predictions effectively. We propose \textbf{GraphGeo}, a multi-agent debate framework using heterogeneous graph neural networks for visual geo-localization. Our approach models diverse debate relationships through typed edges, distinguishing supportive collaboration, competitive argumentation, and knowledge transfer. We introduce a dual-level debate mechanism combining node-level refinement and edge-level argumentation modeling. A cross-level topology refinement strategy enables co-evolution between graph structure and agent representations. Experiments on multiple benchmarks demonstrate GraphGeo significantly outperforms state-of-the-art methods. Our framework transforms cognitive conflicts between agents into enhanced geo-localization accuracy through structured debate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。