提出可理解空间信息的无人机车辆识别系统,支持精准属性检索。
AirSpatialBot: A Spatially-Aware Aerial Agent for Fine-Grained Vehicle Attribute Recognization and Retrieval
- 构建包含206K指令的空域感知数据集AirSpatial,引入3DBB标注
- 通过两阶段训练提升模型对航拍图像的空间理解能力
- 开发能动态规划任务的空中智能体,支持细粒度车辆属性识别
尽管遥感视觉语言模型(VLMs)取得显著进展,但现有模型普遍存在空间理解能力不足的问题,限制了其在真实场景中的应用。为突破遥感领域VLMs的边界,本文聚焦无人机拍摄的车辆图像,提出一个具备空间感知能力的数据集AirSpatial,包含超过206,000条指令,并引入两项新任务:空间定位与空间问答。该数据集是首个提供3DBB标注的遥感定位数据集。为有效利用现有VLM在图像理解方面的优势,我们采用两阶段训练策略:图像理解预训练与空间理解微调。基于此训练好的空间感知VLM,我们构建了空中智能体AirSpatialBot,具备细粒度车辆属性识别与检索能力。通过动态融合任务规划、图像理解、空间理解与任务执行能力,AirSpatialBot可适应多样化的查询需求。实验验证了方法的有效性,揭示了现有VLMs在空间理解上的局限性并提供了重要洞见。模型、代码与数据集将开源。
原文摘要 · Abstract (English)
Despite notable advancements in remote sensing vision-language models (VLMs), existing models often struggle with spatial understanding, limiting their effectiveness in real-world applications. To push the boundaries of VLMs in remote sensing, we specifically address vehicle imagery captured by drones and introduce a spatially-aware dataset AirSpatial, which comprises over 206K instructions and introduces two novel tasks: Spatial Grounding and Spatial Question Answering. It is also the first remote sensing grounding dataset to provide 3DBB. To effectively leverage existing image understanding of VLMs to spatial domains, we adopt a two-stage training strategy comprising Image Understanding Pre-training and Spatial Understanding Fine-tuning. Utilizing this trained spatially-aware VLM, we develop an aerial agent, AirSpatialBot, which is capable of fine-grained vehicle attribute recognition and retrieval. By dynamically integrating task planning, image understanding, spatial understanding, and task execution capabilities, AirSpatialBot adapts to diverse query requirements. Experimental results validate the effectiveness of our approach, revealing the spatial limitations of existing VLMs while providing valuable insights. The model, code, and datasets will be released at https://github.com/VisionXLab/AirSpatialBot
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。