无人机与卫星协同推理,提升复杂环境下的空间理解能力。
AeroVerse-SatAgent: UAV-Satellite Collaborative Spatial Reasoning Inspired by the Dual Visual Pathway Theory of Cognitive Neuroscience

- 借鉴人类视觉双通路理论,融合卫星与无人机多视角信息。
- 在复杂城市环境中,空间关系推理准确率领先现有模型25.91%。
- 适合遥感、智能交通等需要多视角几何推理的场景。
随着航空航天体化智能的快速发展,使无人机自主理解并推理复杂环境变得日益重要。然而,现有基于无人机的空间推理方法存在明显局限:单一视角感知易受遮挡和视角畸变影响,多数视觉语言模型缺乏显式几何建模,依赖语义线索,在视角与尺度变化下推理结果不一致。为此,我们提出SatAgent,一种受人类视觉系统双通路机制启发的无人机-卫星协同空间推理模型。通过联合利用卫星与无人机视角,实现复杂城市环境中的鲁棒、精准推理。首先引入几何感知3D重建编码器,将2D无人机特征提升为显式3D空间表示;其次设计多视角拓扑-语义对齐模块,在统一鸟瞰图(BEV)坐标系中融合跨视角特征;进一步提出多视角一致性损失,促进视角不变表示学习。最后构建了首个大规模无人机-卫星协同多视角空间推理数据集SatAgent-SR130K。实验表明,SatAgent在多样任务上超越最先进的通用基础模型25.91%、专用空间推理模型11.69%,尤其在复杂几何关系推理中表现突出。
原文摘要 · Abstract (English)
With the rapid advancement of aerospace embodied intelligence, enabling Unmanned Aerial Vehicles (UAVs) to autonomously understand and reason about complex environments has become increasingly important. However, existing UAV-based spatial reasoning approaches face critical limitations: single-view perception renders them vulnerable to occlusions and perspective distortions, while most VLMs lack explicit geometric modeling, relying on semantic cues and yielding inconsistent reasoning under viewpoint and scale variations. To address these challenges, we propose SatAgent, a UAV-Satellite collaborative spatial reasoning model inspired by the dual-pathway mechanism of the human visual system. By jointly leveraging satellite and UAV perspectives, SatAgent enables robust, accurate reasoning in complex urban environments. We first introduce a Geometric-Aware 3D Reconstruction Encoder that elevates 2D UAV features into explicit 3D spatial representations. Next, we design a multi-view topology-semantic alignment module integrating cross-view features within a unified BEV coordinate system. We further introduce a multi-view consistency loss encouraging viewpoint-invariant representations. Finally, we construct SatAgent-SR130K, the first large-scale UAV-Satellite collaborative multi-view spatial reasoning dataset. Experiments show SatAgent outperforms state-of-the-art general-purpose foundation models and specialized spatial reasoning models by 25.91\% and 11.69\%, respectively, across diverse tasks, achieving particularly high accuracy in complex geometric relationship reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。