让无人机和地面机器人协作找车,用新基准和智能方法提升搜索效率。
Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method

- 设计协作搜索流程,让视觉语言模型专注理解与决策。
- 在7700个任务中,最高成功率从8.6%提升至55.7%。
- 适合研究多智能体协作、视觉语言模型应用的开发者。
城市环境中无人机与地面机器人协同搜寻目标车辆是一项具身挑战任务,需通过多视角视觉信息联合发现并验证目标。为研究这一未被充分探索的问题,我们提出AGOS-Bench,首个用于评估通用视觉语言模型(VLMs)在无人机-地面机器人协作中融合空中发现与地面验证能力的基准。同时提供配套数据集AGOS-Dataset,由自动化流水线构建,包含7.7千个不同类别与属性目标的搜寻任务,涵盖三个难度等级。针对该任务,我们提出AGOS-Agent——一种无需训练、工具增强的智能体方法。该方法通过明确的“搜索-移交-验证”协作协议,减轻视觉语言模型在复杂动态协调中的负担,仅要求其完成场景理解与决策。在九种VLM上进行的实验表明,AGOS-Agent使其中八种模型整体成功率达更高,且所有模型决策步数均减少。在高难度子集上,Gemini-3.6-Flash的成功率(SR)从8.6%提升至55.7%,路径相似度(SPL)从7.6%升至44.0%。
原文摘要 · Abstract (English)
Air-Ground Object Search (AGOS) in urban environments is a challenging embodied task, which requires an Unmanned Aerial Vehicle (UAV) and an Unmanned Ground Vehicle (UGV) to jointly search for and verify a specified target vehicle from multi-view visual references. To study this underexplored problem, we introduce AGOS-Bench, the first dedicated benchmark for evaluating whether general-purpose Vision-Language Models (VLMs) can integrate aerial discoveries and ground-level verification through UAV-UGV cooperation. We further provide AGOS-Dataset as the companion resource of exemplary trajectories constructed by an automatic pipeline. It consists of 7.7k episodes for searching objects of diverse categories and attributes, spanning three difficulty levels. To address the AGOS task, we propose AGOS-Agent, a training-free and tool-augmented approach. The agentic method relieves VLMs from complex and dynamic coordination via a deliberate search-handoff-verify cooperation protocol, only demanding VLMs for scene understanding and decision-making. Extensive experiments on nine VLMs show that AGOS-Agent improves overall success rate for eight of the nine evaluated backbones while reducing decision steps for all nine. On the hard split, the SR and SPL of Gemini-3.6-Flash increase from 8.6% to 55.7% and from 7.6% to 44.0%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。