arXiv:2511.11025cs.CVcs.AI2025-11AAAI被引 20

首个评估多无人机协作感知的基准,聚焦真实复杂场景下的智能决策。

AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning

  • 构建包含14.6万+问题的多无人机协同感知基准,覆盖四大任务维度。
  • 40个MLLM模型在任务中平均比人类低24.38%,且表现不一致。
  • 适用于多智能体协同视觉、无人机系统与具身推理研究者。

多模态大语言模型(MLLMs)在单智能体视觉任务中表现优异,但评估多智能体协同感知的基准仍匮乏。多无人机系统相比单传感器方案具备更广覆盖、更强鲁棒性与协作能力。现有多图像基准主要针对高质量单智能体图像的简单感知任务,难以评估MLLMs在复杂、第一人称协同场景下的表现,尤其在真实世界感知退化条件下。为此,我们提出AirCopBench,首个面向具身空中协同感知与推理的综合性基准。该基准涵盖14.6万+问题,来自模拟器与真实数据,覆盖场景理解、物体理解、感知评估与协同决策四个维度,共14类任务。通过在感知退化场景下标注协同事件,结合模型、规则与人工生成方法,在严格质量控制下构建大规模问题集。对40个MLLMs的评估显示,协同感知任务存在显著性能差距,最优模型平均落后人类24.38%,且任务间表现不一致。微调实验进一步验证了仿真到现实迁移在空中协同感知与推理中的可行性。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown promise in single-agent vision tasks, yet benchmarks for evaluating multi-agent collaborative perception remain scarce. This gap is critical, as multi-drone systems provide enhanced coverage, robustness, and collaboration compared to single-sensor setups. Existing multi-image benchmarks mainly target basic perception tasks using high-quality single-agent images, thus failing to evaluate MLLMs in more complex, egocentric collaborative scenarios, especially under real-world degraded perception conditions.To address these challenges, we introduce AirCopBench, the first comprehensive benchmark designed to evaluate MLLMs in embodied aerial collaborative perception under challenging perceptual conditions. AirCopBench includes 14.6k+ questions derived from both simulator and real-world data, spanning four key task dimensions: Scene Understanding, Object Understanding, Perception Assessment, and Collaborative Decision, across 14 task types. We construct the benchmark using data from challenging degraded-perception scenarios with annotated collaborative events, generating large-scale questions through model-, rule-, and human-based methods under rigorous quality control. Evaluations on 40 MLLMs show significant performance gaps in collaborative perception tasks, with the best model trailing humans by 24.38% on average and exhibiting inconsistent results across tasks. Fine-tuning experiments further confirm the feasibility of sim-to-real transfer in aerial collaborative perception and reasoning.

多智能体无人机具身智能大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。