构建无人机双认知推理基准,评估模型自知与环境感知能力。
Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs

- 从无人机自身状态与外部环境双视角设计多视图时空推理任务
- 包含上千个问答样本,需精确空间与时间定位,非简单分类
- 验证现有大模型在视角变换、时空定位上仍存显著缺陷
多模态大语言模型在视觉-语言任务中表现强劲,但在无人机场景中的能力仍不充分。现有无人机基准多聚焦于场景理解或事件识别,缺乏对无人机代理所需双重认知能力的联合评估:即在多视角时空上下文中对自身状态和外部环境进行推理。为此,我们提出UAV-DualCog,一个基于双重认知视角的无人机多视图时空推理基准。该基准包含图像与视频任务,同时评估自我状态与环境状态推理,要求超越离散答案预测的空间或时间定位。我们还开发了自动化数据构建管道,基于场景级语义点云生成可扩展数据集,涵盖多样场景、数百个地标和数千个问答样本。大量评估显示,当前多模态大模型在无人机双重认知任务中仍不可靠。自我状态推理、视角变换、精确空间定位及时间区间定位是持续瓶颈。通过思维链/前沿模型和人工基线验证,确认该基准对人类可理解但对现有模型极具挑战性。我们进一步构建了来自非重叠场景的UAV-DualCog-Train,并通过轻量优化探测证实其提供有效结构化监督,表明其不仅是评估基准,也可作为提升多模态大模型无人机智能体的数据资源。
原文摘要 · Abstract (English)
Multimodal large language models have achieved strong performance across diverse vision-language tasks, yet their capabilities in UAV scenarios remain insufficiently explored. Recent UAV-oriented benchmarks have begun to evaluate MLLMs in aerial scenarios, but they typically focus on scene understanding, event recognition, or navigation completion, rather than jointly assessing the dual-cognition capability required for UAV agents: reasoning about both the UAV's own state and the external environment in multiview spatio-temporal contexts. To address this gap, we present UAV-DualCog, a benchmark for aerial multiview spatio-temporal reasoning built on this dual-cognition perspective. UAV-DualCog includes both image and video tasks to jointly evaluate self-state and environment-state reasoning, while requiring spatial or temporal grounding beyond discrete answer prediction. We also develop an automated pipeline that constructs data from scene-level semantic point clouds, yielding a scalable benchmark with diverse scenes, hundreds of landmarks, and thousands of QA samples. Extensive evaluations show that current MLLMs remain far from reliable in UAV dual cognition. Self-state reasoning, viewpoint transformation, precise spatial grounding, and temporal interval localization are persistent bottlenecks, and additional validation with thinking/frontier models and a human baseline confirms that the benchmark is understandable to humans but challenging for existing models. We further construct UAV-DualCog-Train from disjoint scenes and show through a lightweight optimization probe that it provides useful structured supervision, suggesting its value not only as an evaluation benchmark but also as a data resource for advancing MLLM-based UAV agents. Project website and supplementary materials: https://uav-dualcog.lozumi.com
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。