构建大规模语言引导的遥感目标检测数据集,提升细粒度开放世界识别能力。
OS-W2S: An Automatic Labeling Engine for Language-Guided Open-Set Aerial Object Detection
- 基于视觉-语言模型与BERT后处理,实现从词到句的自动标注
- 构建含200万图文对的MI-OAD数据集,规模达同类数据40倍
- 显著提升零样本迁移性能,推动遥感图像语义理解发展
近年来,语言引导的开放集遥感目标检测因其更贴近真实应用需求而受到关注。然而,受限于数据集规模,现有方法多聚焦词汇级描述,难以满足细粒度开放世界检测需求。为此,我们构建了一个大规模语言引导的开放集遥感目标检测数据集,涵盖词、短语到句子三个层级的语言指导。依托开源大视觉-语言模型,结合图像操作预处理与BERT后处理,提出OS-W2S标签引擎,实现对航拍图像多样场景的自动标注。利用该引擎,我们扩展了现有遥感检测数据集并构建新基准数据集MI-OAD,弥补当前遥感定位数据的不足,支持有效的语言引导开放集检测。MI-OAD包含163,023张图像和200万组图像-文本对,约为现有同类数据的40倍。为验证其有效性与质量,我们评估了三项代表性任务:在语言引导的开放集遥感检测中,基于MI-OAD训练使Grounding DINO在零样本迁移下,句子输入时的AP$_{50}$提升31.1,Recall@10提升34.7;此外,使用MI-OAD进行预训练,在多个现有开放词汇遥感检测与遥感视觉定位基准上达到最先进性能,验证了数据集的有效性与标注高质量。更多详情请见https://github.com/GT-Wei/MI-OAD。
原文摘要 · Abstract (English)
In recent years, language-guided open-set aerial object detection has gained significant attention due to its better alignment with real-world application needs. However, due to limited datasets, most existing language-guided methods primarily focus on vocabulary-level descriptions, which fail to meet the demands of fine-grained open-world detection. To address this limitation, we propose constructing a large-scale language-guided open-set aerial detection dataset, encompassing three levels of language guidance: from words to phrases, and ultimately to sentences. Centered around an open-source large vision-language model and integrating image-operation-based preprocessing with BERT-based postprocessing, we present the OS-W2S Label Engine, an automatic annotation pipeline capable of handling diverse scene annotations for aerial images. Using this label engine, we expand existing aerial detection datasets with rich textual annotations and construct a novel benchmark dataset, called MI-OAD, addressing the limitations of current remote sensing grounding data and enabling effective language-guided open-set aerial detection. Specifically, MI-OAD contains 163,023 images and 2 million image-caption pairs, approximately 40 times larger than comparable datasets. To demonstrate the effectiveness and quality of MI-OAD, we evaluate three representative tasks. On language-guided open-set aerial detection, training on MI-OAD lifts Grounding DINO by +31.1 AP$_{50}$ and +34.7 Recall@10 with sentence-level inputs under zero-shot transfer. Moreover, using MI-OAD for pre-training yields state-of-the-art performance on multiple existing open-vocabulary aerial detection and remote sensing visual grounding benchmarks, validating both the effectiveness of the dataset and the high quality of its OS-W2S annotations. More details are available at https://github.com/GT-Wei/MI-OAD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。