arXiv:2505.08765cs.CVcs.AI2025-05被引 1

提出首个城市无人机视觉搜物基准与智能搜索方法,提升自主寻物效率。

Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic Methodology

  • 用多模态大模型构建感知-推理-规划三重认知机制
  • 在6类物体2420个任务上实现成功率提升37.69%、搜索效率提升28.96%
  • 适合研究无人机自主导航与具身智能的开发者参考

城市环境中的空中视觉目标搜索(AVOS)要求无人机在无外部引导下,利用视觉和文本线索自主搜寻并识别目标。现有方法在复杂城市环境中表现不佳,主要受限于冗余语义处理、相似物体区分困难以及探索与利用的权衡问题。为填补这一空白,我们提出首个面向城市场景的自主搜物基准数据集CityAVOS,包含2420个任务,覆盖六类常见城市物体,难度分层明确,支持对无人机智能体搜索能力的全面评估。针对该任务,我们提出PRPSearcher(感知-推理-规划搜索器),一种基于多模态大语言模型的新型智能体方法,模拟人类三层认知过程:构建以目标为中心的动态语义地图增强空间感知,基于语义吸引力值建立3D认知地图用于目标推理,以及基于不确定性值的3D探索地图实现平衡的探索-利用策略。同时引入去噪机制降低相似物体干扰,并设计灵感促进思维(IPT)提示机制实现自适应动作规划。在CityAVOS上的实验表明,PRPSearcher在成功率(+37.69%)、搜索路径长度(+28.96% SPL)、平均搜寻步数(-30.69% MSS)和行动次数(-46.40% NE)上均显著优于现有基线。尽管表现优异,但与人类水平仍有差距,凸显了在语义推理与空间探索方面仍需突破。本工作为具身目标搜索的未来发展奠定了基础。数据集与代码已公开于https://anonymous.4open.science/r/CityAVOS-3DF8。

原文摘要 · Abstract (English)

Aerial Visual Object Search (AVOS) tasks in urban environments require Unmanned Aerial Vehicles (UAVs) to autonomously search for and identify target objects using visual and textual cues without external guidance. Existing approaches struggle in complex urban environments due to redundant semantic processing, similar object distinction, and the exploration-exploitation dilemma. To bridge this gap and support the AVOS task, we introduce CityAVOS, the first benchmark dataset for autonomous search of common urban objects. This dataset comprises 2,420 tasks across six object categories with varying difficulty levels, enabling comprehensive evaluation of UAV agents' search capabilities. To solve the AVOS tasks, we also propose PRPSearcher (Perception-Reasoning-Planning Searcher), a novel agentic method powered by multi-modal large language models (MLLMs) that mimics human three-tier cognition. Specifically, PRPSearcher constructs three specialized maps: an object-centric dynamic semantic map enhancing spatial perception, a 3D cognitive map based on semantic attraction values for target reasoning, and a 3D uncertainty map for balanced exploration-exploitation search. Also, our approach incorporates a denoising mechanism to mitigate interference from similar objects and utilizes an Inspiration Promote Thought (IPT) prompting mechanism for adaptive action planning. Experimental results on CityAVOS demonstrate that PRPSearcher surpasses existing baselines in both success rate and search efficiency (on average: +37.69% SR, +28.96% SPL, -30.69% MSS, and -46.40% NE). While promising, the performance gap compared to humans highlights the need for better semantic reasoning and spatial exploration capabilities in AVOS tasks. This work establishes a foundation for future advances in embodied target search. Dataset and source code are available at https://anonymous.4open.science/r/CityAVOS-3DF8.

无人机视觉搜索多模态智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。