arXiv:2505.18229cs.ROcs.AI2025-05被引 15

构建首个面向无人机智能体的标准化评估基准,推动自主飞行任务研究

BEDI: A Comprehensive Benchmark for Evaluating Embodied Agents on UAVs

  • 提出动态任务链范式,将复杂飞行任务拆解为可测量的子任务
  • 涵盖六项核心能力评估,覆盖感知、控制、规划等关键环节
  • 支持虚实混合场景与开放接口,适合算法研究与系统开发人员使用

随着低空遥感与视觉语言模型(VLMs)的快速发展,基于无人机(UAV)的具身智能体在自主任务中展现出巨大潜力。然而,当前对无人机具身智能体(UAV-EAs)的评估仍受限于缺乏标准化基准、多样化测试场景及开放系统接口。为此,本文提出BEDI(具身无人机智能评估基准),一个系统化、标准化的评估框架。我们引入基于感知-决策-行动循环的新型动态任务链范式,将复杂无人机任务分解为标准化、可度量的子任务。在此基础上,设计涵盖六项核心子技能的统一评估框架:语义感知、空间感知、运动控制、工具使用、任务规划与动作生成。进一步构建融合虚拟与真实场景的混合测试平台,实现跨环境的全面评估。平台提供开放标准接口,支持任务定制与场景扩展,提升评估灵活性与可扩展性。通过对多个前沿VLMs的实证评估,揭示其在具身无人机任务中的局限性,凸显BEDI在推动具身智能研究与模型优化中的关键作用。通过填补该领域系统化评估的空白,BEDI实现了客观模型对比,为未来研究奠定坚实基础。基准已开源:https://github.com/lostwolves/BEDI。

原文摘要 · Abstract (English)

With the rapid advancement of low-altitude remote sensing and Vision-Language Models (VLMs), Embodied Agents based on Unmanned Aerial Vehicles (UAVs) have shown significant potential in autonomous tasks. However, current evaluation methods for UAV-Embodied Agents (UAV-EAs) remain constrained by the lack of standardized benchmarks, diverse testing scenarios and open system interfaces. To address these challenges, we propose BEDI (Benchmark for Embodied Drone Intelligence), a systematic and standardized benchmark designed for evaluating UAV-EAs. Specifically, we introduce a novel Dynamic Chain-of-Embodied-Task paradigm based on the perception-decision-action loop, which decomposes complex UAV tasks into standardized, measurable subtasks. Building on this paradigm, we design a unified evaluation framework encompassing six core sub-skills: semantic perception, spatial perception, motion control, tool utilization, task planning and action generation. Furthermore, we develop a hybrid testing platform that incorporates a wide range of both virtual and real-world scenarios, enabling a comprehensive evaluation of UAV-EAs across diverse contexts. The platform also offers open and standardized interfaces, allowing researchers to customize tasks and extend scenarios, thereby enhancing flexibility and scalability in the evaluation process. Finally, through empirical evaluations of several state-of-the-art (SOTA) VLMs, we reveal their limitations in embodied UAV tasks, underscoring the critical role of the BEDI benchmark in advancing embodied intelligence research and model optimization. By filling the gap in systematic and standardized evaluation within this field, BEDI facilitates objective model comparison and lays a robust foundation for future development in this field. Our benchmark is now publicly available at https://github.com/lostwolves/BEDI.

具身智能无人机评估基准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。