针对无人机图像开放词汇检测难题,构建新数据集并提出轻量增强模块。
Light-Weight Cross-Modal Enhancement Method with Benchmark Construction for UAV-based Open-Vocabulary Object Detection
- 设计无人机标注引擎,生成大规模高质量图文数据
- 在VisDrone上零样本检测提升5.3 mAP,参数与计算量更低
- 适合无人机视觉、跨域检测研究者使用
开放词汇目标检测(OVD)在应用于无人机影像时,因与地面数据集存在领域差异导致性能严重下降。为此,我们提出一套面向无人机的完整解决方案,结合数据集构建与模型创新。首先,设计精炼的UAV-Label Engine,高效解决标注冗余、不一致与模糊问题,实现大规模无人机数据集生成。基于此引擎,构建两个新基准:包含超过240万实例、覆盖1800+类别的UAVDE-2M,以及提供丰富图文对的UAVCAP-15K,用于视觉语言预训练。其次,提出交叉注意力门控增强(CAGE)模块,一种轻量级双路径融合结构,集成交叉注意力、自适应门控与全局FiLM调制,实现鲁棒的文本-视觉对齐。将CAGE嵌入YOLO-World-v2框架后,显著提升精度与效率,在VisDrone上零样本检测提升5.3 mAP,同时减少参数与GFLOPs,且在SIMD数据集上展现强跨域泛化能力。大量实验与真实无人机部署验证了该方案的有效性与实用性。
原文摘要 · Abstract (English)
Open-Vocabulary Object Detection (OVD) faces severe performance degradation when applied to UAV imagery due to the domain gap from ground-level datasets. To address this challenge, we propose a complete UAV-oriented solution that combines both dataset construction and model innovation. First, we design a refined UAV-Label Engine, which efficiently resolves annotation redundancy, inconsistency, and ambiguity, enabling the generation of largescale UAV datasets. Based on this engine, we construct two new benchmarks: UAVDE-2M, with over 2.4M instances across 1,800+ categories, and UAVCAP-15K, providing rich image-text pairs for vision-language pretraining. Second, we introduce the Cross-Attention Gated Enhancement (CAGE) module, a lightweight dual-path fusion design that integrates cross-attention, adaptive gating, and global FiLM modulation for robust textvision alignment. By embedding CAGE into the YOLO-World-v2 framework, our method achieves significant gains in both accuracy and efficiency, notably improving zero-shot detection on VisDrone by +5.3 mAP while reducing parameters and GFLOPs, and demonstrating strong cross-domain generalization on SIMD. Extensive experiments and real-world UAV deployment confirm the effectiveness and practicality of our proposed solution for UAV-based OVD
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。