首次评估开放词汇检测在航拍图像上的迁移能力,发现性能严重下降。
Do Open-Vocabulary Detectors Transfer to Aerial Imagery? A Comparative Evaluation
- 构建航拍数据集LA-80C,测试五种SOTA模型在零样本条件下的表现
- 最佳模型F1仅27.6%,假阳性率达69%,词汇量过大会导致语义混淆
- 提示工程无效,模型对成像条件敏感,适合研究航拍场景的视觉语言模型
开放词汇目标检测(OVD)通过视觉-语言模型实现新类别的零样本识别,在自然图像上表现优异,但其在航拍图像中的可迁移性尚未被探索。本文首次系统性地在包含3,592张图像、80个类别的LAE-80C航拍数据集上,严格零样本条件下评估了五种SOTA OVD模型。实验采用全局、理想和单类别推理模式,分离语义混淆与视觉定位问题。结果表明:域迁移严重失败,最佳模型OWLv2的F1仅为27.6%,假阳性率高达69%。将词汇量从80缩减至3.2类后性能提升15倍,证明语义混淆是主要瓶颈。领域特定前缀和同义词扩展等提示工程策略未带来显著提升。不同数据集间性能差异巨大(DIOR F1=0.53,FAIR1M F1=0.12),暴露模型对成像条件的脆弱性。研究确立了航拍OVD的基准预期,强调需发展域自适应方法。
原文摘要 · Abstract (English)
Open-vocabulary object detection (OVD) enables zero-shot recognition of novel categories through vision-language models, achieving strong performance on natural images. However, transferability to aerial imagery remains unexplored. We present the first systematic benchmark evaluating five state-of-the-art OVD models on the LAE-80C aerial dataset (3,592 images, 80 categories) under strict zero-shot conditions. Our experimental protocol isolates semantic confusion from visual localization through Global, Oracle, and Single-Category inference modes. Results reveal severe domain transfer failure: the best model (OWLv2) achieves only 27.6% F1-score with 69% false positive rate. Critically, reducing vocabulary size from 80 to 3.2 classes yields 15x improvement, demonstrating that semantic confusion is the primary bottleneck. Prompt engineering strategies such as domain-specific prefixing and synonym expansion, fail to provide meaningful performance gains. Performance varies dramatically across datasets (F1: 0.53 on DIOR, 0.12 on FAIR1M), exposing brittleness to imaging conditions. These findings establish baseline expectations and highlight the need for domain-adaptive approaches in aerial OVD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。