arXiv:2607.12477cs.CV2026-07被引 3

提出SIS-Bench基准,评估无人机在空间中的自我意识与认知能力。

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

论文配图:Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence
图 1 · 摘自论文原文
  • 构建双维度(空间与自我)三层次(感知-记忆-推理)评估框架。
  • 4856个问答对揭示当前模型在自我意识上表现远弱于空间理解。
  • 引入运动感知表示,提升感知、记忆及决策性能,适合机器人研发者参考。

自主无人机系统越来越多依赖多模态大语言模型(MLLMs)在复杂真实环境中运行。此类具身场景不仅需要理解周围空间,还需保持对自身状态的一致表征。然而,现有无人机相关方法与基准仍以环境为中心,主要关注空间理解任务,而对智能体的自我意识缺乏显式建模。为弥补这一空白,我们提出SIS-Bench,一个基于统一“自我在空间中”范式的无人机具身空间智能评估基准。SIS-Bench从空间与自我两个互补维度出发,建立感知、记忆、推理三层结构,包含4,856个问答对,覆盖13项任务,源自1,646段真实无人机视频,经任务条件化生成与专家验证。大量实验表明,当前MLLM在建模动态与以代理为中心过程方面存在根本性局限,表现为空间认知与自我意识间的显著失衡,以及认知层级上的渐进性能下降。受此启发,我们进一步探索一种融合光流与视觉特征的运动感知表示。实验显示,建模代理运动可一致提升感知与记忆表现,不仅改善空间认知,也增强自我意识,并泛化至下游无人机决策任务。结果强调了自我意识对推进具身空间智能的重要性,提供了新基准与运动感知自在我空间建模的实证支持。

原文摘要 · Abstract (English)

Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also maintaining a coherent representation of the agent itself. However, existing UAV-oriented approaches and benchmarks remain largely environment-centric, primarily focusing on spatial understanding tasks, with the agent's self-awareness remaining implicit. To address this gap, we introduce SIS-Bench, a benchmark for evaluating embodied spatial intelligence in UAV scenarios under a unified self-in-space formulation. SIS-Bench organizes evaluation along two complementary dimensions, space and self, and a three-level hierarchy of perception, memory, and reasoning. It contains 4,856 question--answer pairs across 13 tasks derived from 1,646 real-world UAV videos through a task-conditioned construction pipeline with expert verification. Extensive evaluations reveal that current MLLMs exhibit fundamental limitations in modeling dynamic and agent-centered processes. In particular, we observe a clear imbalance between spatial cognition and self-awareness, as well as a progressive performance degradation across cognitive levels. Motivated by these findings, we further explore a motion-aware representation that incorporates self-related dynamics through optical flow and visual feature fusion. Experimental results show that modeling agent motion consistently improves perception and memory performance, not only in spatial cognition but also in self-awareness, and generalizes to downstream UAV decision-making tasks. Our results highlight the importance of self-awareness for advancing embodied spatial intelligence, and provide both a new benchmark and empirical evidence for motion-aware self-in-space modeling.

无人机自我意识空间认知具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。