arXiv:2607.10180cs.ROcs.AI2026-07被引 1

首个面向无人机主动感知的基准,连接视觉语言与飞行控制

ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

论文配图:ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception
图 1 · 摘自论文原文
  • 分三层任务构建无人机主动感知框架
  • 真实与仿真环境采集数据,验证模型表现受限于规划能力
  • 适合研究多模态智能体与无人系统交互的学者

我们提出 ActiveFly-Bench,首个连接虚拟世界推理与物理世界交互的无人机具身感知基准。该基准将主动感知分解为三个层级任务:空中具身问答(Air-EQA)、观测行为规划(OBP)和细粒度语言引导无人机控制(FLUC),明确关联高层任务理解、行为规划与底层控制。数据集涵盖真实世界与仿真室外环境,用于训练与评估。我们进一步开发了 ActiveFly,一个闭环无人机智能体,融合视觉-语言推理与细粒度控制,并部署于真实无人机平台。对代表性 VLM 与 VLA 模型的实验表明,当前无人机智能体在行为规划、视角调整及任务鲁棒完成方面仍存在明显短板。这些结果确立了 ActiveFly-Bench 作为具身空中智能新测试平台的地位。

原文摘要 · Abstract (English)

We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. The benchmark decomposes active perception into three hierarchical tasks: Aerial Embodied Question Answering (Air-EQA), Observation Behavior Planning (OBP), and Fine-grained Language-guided UAV Control (FLUC), explicitly connecting high-level task understanding, behavior planning, and low-level control. The datasets are collected from both real-world and simulated outdoor environments for training and evaluation. We further develop ActiveFly, a closed-loop UAV agent that integrates visual-language reasoning with fine-grained control, and deploy it on a physical UAV platform. Experiments with representative VLMs and VLA models show that current UAV agents still struggle with behavior planning, viewpoint adjustment, and robust task completion in active perception. These results establish ActiveFly-Bench as a new testbed for embodied aerial intelligence.

无人机感知视觉语言具身智能基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。