通过分阶段推理提升无人机图像中目标定位精度
GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery

- 先生成可靠候选目标,再用图注意力融合语言与空间关系
- 在AerialVG和AerialSense上分别达到67.31%和80.34%准确率
- 适合高密度复杂场景下的视觉定位任务
无人机图像中的视觉定位旨在根据自然语言描述定位复杂鸟瞰场景中的目标物体。然而,大量小而密集、外观相似的物体导致高视觉冗余,重复的局部结构引发强拓扑歧义。现有方法多关注视觉-语言特征对齐或密集上下文交互,难以区分细微的实例差异并有效利用空间拓扑结构,导致在高密度场景中定位不准。为此,我们提出GrabVG,一种受人类视觉搜索启发的新框架。该框架将定位过程分为两个阶段:前注意假设搜索与图注意力特征绑定。首先通过知识蒸馏引导的提案生成与文本感知的假设过滤,大幅减少背景干扰和语义错配;随后将候选目标组织为稀疏图,通过图注意力联合传播语言引导的内部视觉线索与实例间拓扑关系,实现高效空间推理与精准定位。在AerialVG和AerialSense上的实验表明,GrabVG在准确率与速度之间取得良好平衡,分别达到67.31%和80.34%的[email protected],优于对应基线10.55和8.76个百分点。
原文摘要 · Abstract (English)
Visual grounding in Unmanned Aerial Vehicle (UAV) imagery aims to localize a target object in complex bird's-eye-view scenes according to a natural language description. However, the abundance of small, densely distributed, and visually similar objects creates high visual redundancy, while repetitive local configurations give rise to strong topological ambiguity. Existing approaches mainly focus on visual--language feature alignment or dense contextual interaction, yet they struggle to distinguish subtle inter-instance differences and effectively exploit spatial topological structures, leading to inaccurate grounding in highly crowded scenarios. To address these challenges, we propose $\textbf{GrabVG}$, a novel visual grounding framework inspired by human visual search. GrabVG explicitly decomposes grounding into two sequential stages: $\textit{preattentive hypothesis search}$ and $\textit{graph-attentive feature binding}$. Specifically, we first generate a compact set of reliable object hypotheses through distillation-guided proposal induction and text-aware hypothesis filtering, substantially reducing background distractions and semantic mismatches. These hypotheses are then organized into a sparse graph, where language-guided intra-instance visual cues and inter-instance topological relationships are jointly bound and propagated via graph attention, enabling efficient spatial reasoning and accurate target localization. Extensive experiments on AerialVG and AerialSense show that GrabVG achieves a favorable accuracy--speed trade-off, reaching 67.31$\%$ and 80.34$\%$ [email protected] and outperforming the corresponding baselines by 10.55 and 8.76 percentage points, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。