arXiv:2604.20543cs.CV2026-04

为航拍图像中的指代检测构建新数据集与高效算法

RefAerial: A Benchmark and Approach for Referring Detection in Aerial Images

论文配图:RefAerial: A Benchmark and Approach for Referring Detection in Aerial Images
图 1 · 摘自论文原文
  • 设计航拍图像指代检测数据集RefAerial,包含多样低占比目标
  • 提出SCS框架,在航拍数据上显著提升指代定位准确率
  • 适合研究遥感视觉、跨模态定位的学者参考

指代检测旨在定位自然语言描述的目标,近年来受到广泛关注。然而,现有数据集多局限于地面图像,目标居中且场景较小。本文引入大规模挑战性航拍图像指代检测数据集RefAerial,具备四大特征:(1)目标与场景比例低且多样,(2)目标与干扰物数量众多,(3)描述复杂精细,(4)航拍视角场景丰富多样。我们还开发了人机协同的指代对生成与标注引擎(REA-Engine),实现半自动化高效标注。此外,发现现有地面指代检测方法在本航拍数据集上性能严重下降,根源在于图像内或跨图像的尺度差异。为此,我们提出新型尺度全面敏感(SCS)框架,包含粒度混合(MoG)注意力机制和两阶段全面到敏感(CtS)解码策略。其中,粒度混合注意力用于尺度全面的目标理解,两阶段解码策略实现从粗到细的指代目标定位。最终,所提SCS框架在航拍指代检测数据集上表现优异,并在传统地面数据集上也取得显著性能提升。

原文摘要 · Abstract (English)

Referring detection refers to locate the target referred by natural languages, which has recently attracted growing research interests. However, existing datasets are limited to ground images with large object centered in relative small scenes. This paper introduces a large-scale challenging dataset for referring detection in aerial images, termed as RefAerial. It distinguishes from conventional ground referring detection datasets by 4 characteristics: (1) low but diverse object-to-scene ratios, (2) numerous targets and distractors, (3)complex and fine-grained referring descriptions, (4) diverse and broad scenes in the aerial view. We also develop a human-in-the-loop referring expansion and annotation engine (REA-Engine) for efficient semi-automated referring pair annotation. Besides, we observe that existing ground referring detection approaches exhibiting serious performance degradation on our aerial dataset since the intrinsic scale variety issue within or across aerial images. Therefore, we further propose a novel scale-comprehensive and sensitive (SCS) framework for referring detection in aerial images. It consists of a mixture-of-granularity (MoG) attention and a two-stage comprehensive-to-sensitive (CtS) decoding strategy. Specifically, the mixture-of-granularity attention is developed for scale-comprehensive target understanding. In addition, the two-stage comprehensive-to-sensitive decoding strategy is designed for coarse-to-fine referring target decoding. Eventually, the proposed SCS framework achieves remarkable performance on our aerial referring detection dataset and even promising performance boost on conventional ground referring detection datasets.

指代检测航拍图像多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。