arXiv:2604.04017cs.CL2026-04被引 3

构建首个融合视觉推理与多跳验证的地理定位基准,评估智能体工具使用能力。

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces

  • 设计双层级地理定位任务,结合碎片化视觉线索与多跳网络验证
  • 引入专家标注的推理轨迹,支持对决策路径的细粒度分析
  • 提出GATE工作流,证明合理工具规划比增加调用次数更有效

深度研究智能体通过多步工具调用整合零散证据。现有文本类基准如BrowseComp仅限文本交互,而多数多模态基准未要求同时处理弱视觉线索和多跳验证。地理定位天然适合作为测试场景,因其答案需融合多个模糊视觉线索,并通过开放网络证据进行验证。为此,我们提出GeoBrowse,一个结合视觉推理与知识密集型多跳查询的地理定位基准。Level 1测试从图像中提取并组合碎片化视觉线索;Level 2通过引入长尾知识和实体混淆,提升查询难度。为支持评估,我们提供包含五种思维-图像工具和四种知识密集型工具的智能体工作流GATE,以及基于可验证证据的专家标注步骤轨迹,实现轨迹级分析。实验表明,GATE优于直接推理和开源智能体,说明无工具、仅搜索或仅图像的方案均不足。性能提升源于一致且层级匹配的工具使用计划,而非更多调用次数,其能更可靠地抵达关键证据节点,减少整合错误。代码与数据集已公开于https://github.com/ornamentt/GeoBrowse。

原文摘要 · Abstract (English)

Deep research agents integrate fragmented evidence through multi-step tool use. BrowseComp offers a text-only testbed for such agents, but existing multimodal benchmarks rarely require both weak visual cues composition and BrowseComp-style multi-hop verification. Geolocation is a natural testbed because answers depend on combining multiple ambiguous visual cues and validating them with open-web evidence. Thus, we introduce GeoBrowse, a geolocation benchmark that combines visual reasoning with knowledge-intensive multi-hop queries. Level 1 tests extracting and composing fragmented visual cues, and Level 2 increases query difficulty by injecting long-tail knowledge and obfuscating key entities. To support evaluation, we provide an agentic workflow GATE with five think-with-image tools and four knowledge-intensive tools, and release expert-annotated stepwise traces grounded in verifiable evidence for trajectory-level analysis. Experiments show that GATE outperforms direct inference and open-source agents, indicating that no-tool, search-only or image-only setups are insufficient. Gains come from coherent, level-specific tool-use plans rather than more tool calls, as they more reliably reach annotated key evidence steps and make fewer errors when integrating into the final decision. The GeoBrowse bernchmark and codes are provided in https://github.com/ornamentt/GeoBrowse

地理定位智能体多跳推理视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。