arXiv:2605.01371cs.RO2026-05被引 3

首个面向智能无人机搜救的综合性基准,推动真实场景下的自主搜救研究。

ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue

论文配图:ESARBench: A Benchmark for Agentic UAV Embodied Search and Rescue
图 1 · 摘自论文原文
  • 提出新型沉浸式搜救任务ESAR,要求无人机自主探索与推理
  • 构建4个基于真实地理数据的高保真仿真环境,支持动态天气与线索随机分布
  • 提供600个真实救援案例任务和多维评估指标,适合多模态大模型研究者使用

多模态大语言模型(MLLM)的发展使无人机在空间推理、语义理解与复杂决策方面具备强大能力,非常适合用于无人机搜救(SAR)。然而,现有研究仍以传统视觉与路径规划方法为主,缺乏统一的沉浸式智能体评估基准。为此,我们首次提出“沉浸式搜救”(Embodied Search and Rescue, ESAR)任务,要求空中智能体在复杂环境中自主探索、识别救援线索并推理幸存者位置以做出合理决策。同时,我们构建了首个全面的基准ESARBench,利用Unreal Engine 5与AirSim,基于真实地理信息系统(GIS)数据创建了4个大规模高保真开放环境,确保场景高度写实。为真实模拟救援操作,基准引入动态变量:天气、昼夜变化及随机线索布置。我们还构建了一个包含600个任务的数据集,源于真实救援案例,并提出一套稳健的评估指标。通过测试从传统启发式算法到先进地面与空中MLLM-based ObjectNav智能体的多种基线,实验揭示了空间记忆、空中适应性以及搜索效率与飞行安全之间的权衡等关键挑战。我们期望ESARBench能成为推动沉浸式搜救领域发展的有力资源。源码与项目页:https://4amgodvzx.github.io/ESAR.github.io。

原文摘要 · Abstract (English)

The rapid advancement of Multimodal Large Language Models (MLLMs) has empowered Unmanned Aerial Vehicle (UAV) with exceptional capabilities in spatial reasoning, semantic understanding, and complex decision-making, making them inherently suited for UAV Search and Rescue (SAR). However, existing UAV SAR research is dominated by traditional vision and path-planning methods and lacks a comprehensive and unified benchmark for embodied agents. To bridge this gap, we first propose the novel task of \textbf{Embodied Search and Rescue (ESAR)}, which requires aerial agents to autonomously explore complex environments, identify rescue clues, and reason about victim locations to execute informed decision-making. Additionally, we present \textbf{ESARBench}, the first comprehensive benchmark designed to evaluate MLLM-driven UAV agents in highly realistic SAR scenarios. Leveraging Unreal Engine 5 and AirSim, we construct four high-fidelity, large-scale open environments mapped directly from real-world Geographic Information System (GIS) data to ensure photorealistic landscapes. To rigorously simulate actual rescue operations, our benchmark incorporates dynamic variables including weather conditions, time of day, and stochastic clue placement. Furthermore, we create a dataset of 600 tasks modeled after real-world rescue cases and propose a robust set of evaluation metrics. We evaluate diverse baselines, ranging from traditional heuristics to advanced ground and aerial MLLM-based ObjectNav agents. Experimental results highlight the challenges in ESAR, revealing critical bottlenecks in spatial memory, aerial adaptation, and the trade-off between search efficiency and flight safety. We hope ESARBench serves as a valuable resource to advance research on Embodied Search and Rescue domain. Source code and project page: https://4amgodvzx.github.io/ESAR.github.io.

无人机搜救大模型仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。