arXiv:2607.22014cs.AIcs.CL2026-07

测试大模型在空中任务中的零样本推理能力,发现性能远低于人类。

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

论文配图:Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents
图 1 · 摘自论文原文
  • 构建空中3D环境任务基准MissionBench,评估模型零样本规划与导航能力。
  • 最强模型仅完成不足35%任务,人类达84.4%,凸显多步推理难度。
  • 模型规模越大,零样本表现越好,提示通用能力可随规模提升。

多模态大语言模型正成为具身智能体的核心推理模块,但通用模型能否仅凭单一高层指令完成长周期具身任务仍不明确。本文提出MissionBench,一个面向空中3D环境的使命级评估基准,包含120个任务,覆盖五个模拟3D环境和四类任务。智能体需仅基于自我中心观测与动作历史自主规划、导航并报告结果,无需针对空域任务微调。在22个开源与闭源的MLLM中,最强模型成功完成的任务不足35%,而人类达到84.4%,表明多步具身任务极具挑战。尽管不同模型家族表现差异显著,但规模扩增带来性能提升,说明更大通用模型具备更强零样本具身能力。分析显示,任务完成需超越空间感知,整合多步规划与自适应推理。这推动闭环评估,并揭示具身智能规模化改进的潜力与风险。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level instruction. We introduce MissionBench, a benchmark for mission-level evaluation of MLLMs in aerial 3D environments. It comprises 120 missions across five simulated 3D environments and four task families. Agents must autonomously plan, navigate, and report outcomes using only egocentric observations and its action history, without aerial-specific fine-tuning. Across 22 open- and closed-source MLLMs, the strongest model succeeds on fewer than 35% of missions compared to 84.4% human performance, highlighting the difficulty of multi-step embodied tasks. Despite large variations between model families, we observe gains from scaling, indicating that larger general-purpose models possess stronger zero-shot embodied capabilities. Our analysis shows that mission-level competence requires coordinating multiple capabilities beyond spatial perception, including multi-step planning and adaptive reasoning. This motivates closed-loop evaluation and highlights both the promise and risk of scaling-driven improvements for embodied AI.

具身智能多模态大模型零样本任务评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。