arXiv:2504.07836cs.CVcs.AI2025-04ICCV被引 15

构建首个航拍视觉定位数据集,强调空间关系推理。

AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations

  • 设计分层交叉注意力与关系感知模块,聚焦目标区域与空间关系。
  • 数据集含5000张航拍图、5万条描述、10.3万目标,支持多对象定位。
  • 适合关注遥感图像理解、空间推理的科研与工程人员。

视觉定位(VG)旨在根据自然语言描述在图像中定位目标物体。本文提出面向航拍视角的新任务——航拍视觉定位(AerialVG),相比传统视觉定位,其面临新挑战:仅靠外观难以区分多个视觉相似物体,需强化位置关系理解;且现有模型在高分辨率航拍图像上表现不佳。为此,我们构建首个航拍视觉定位数据集AerialVG,包含5,000张真实航拍图像、50,000条人工标注描述和103,000个目标。每条标注均包含多个目标及其相对空间关系,要求模型具备综合空间推理能力。同时,提出专用于AerialVG任务的创新模型,采用分层交叉注意力机制聚焦目标区域,设计关系感知定位模块以推断位置关系。实验验证了数据集与方法的有效性,凸显空间推理在航拍视觉定位中的关键作用。代码与数据集将公开。

原文摘要 · Abstract (English)

Visual grounding (VG) aims to localize target objects in an image based on natural language descriptions. In this paper, we propose AerialVG, a new task focusing on visual grounding from aerial views. Compared to traditional VG, AerialVG poses new challenges, \emph{e.g.}, appearance-based grounding is insufficient to distinguish among multiple visually similar objects, and positional relations should be emphasized. Besides, existing VG models struggle when applied to aerial imagery, where high-resolution images cause significant difficulties. To address these challenges, we introduce the first AerialVG dataset, consisting of 5K real-world aerial images, 50K manually annotated descriptions, and 103K objects. Particularly, each annotation in AerialVG dataset contains multiple target objects annotated with relative spatial relations, requiring models to perform comprehensive spatial reasoning. Furthermore, we propose an innovative model especially for the AerialVG task, where a Hierarchical Cross-Attention is devised to focus on target regions, and a Relation-Aware Grounding module is designed to infer positional relations. Experimental results validate the effectiveness of our dataset and method, highlighting the importance of spatial reasoning in aerial visual grounding. The code and dataset will be released.

视觉定位航拍图像空间推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。