arXiv:2507.13359cs.CV2025-07综述被引 20

让无人机看图识别未知物体,靠自然语言描述实现开放词汇检测。

Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives

  • 利用CLIP等跨模态模型,将文本描述与图像对齐实现开放词汇识别。
  • 系统梳理了无人机影像中开放词汇检测的方法与数据集,发现小目标和遮挡是主要挑战。
  • 适合想进入无人机视觉或开放词汇检测领域的研究者参考。

由于广泛应用,航空图像目标检测长期是计算机视觉的热点。近年来,无人机(UAV)技术的发展进一步推动该领域迈向新高度,催生更广泛的应用需求。然而,传统无人机航拍目标检测方法主要聚焦于预定义类别,严重限制了适用性。跨模态文本-图像对齐(如CLIP)的出现克服了这一局限,实现了开放词汇目标检测(OVOD),可通过自然语言描述识别未见过的物体,显著提升无人机在航拍场景理解中的智能与自主性。本文对无人机航拍场景中的开放词汇目标检测进行了全面综述。首先将OVOD的核心原理与无人机视觉的独特特征对齐,为专项讨论奠定基础。在此基础上,构建了系统化的分类体系,对现有航空影像OVOD方法进行分类,并全面概述相关数据集。该结构化综述使我们能够深入剖析该领域关键挑战与开放问题。最后,基于分析提出有前景的未来研究方向与应用前景。本综述旨在为初学者与资深研究者提供清晰路线图与宝贵参考,推动该快速演进领域的发展。相关工作持续更新:https://github.com/zhouyang2002/OVOD-in-UVA-imagery

原文摘要 · Abstract (English)

Due to its extensive applications, aerial image object detection has long been a hot topic in computer vision. In recent years, advancements in Unmanned Aerial Vehicles (UAV) technology have further propelled this field to new heights, giving rise to a broader range of application requirements. However, traditional UAV aerial object detection methods primarily focus on detecting predefined categories, which significantly limits their applicability. The advent of cross-modal text-image alignment (e.g., CLIP) has overcome this limitation, enabling open-vocabulary object detection (OVOD), which can identify previously unseen objects through natural language descriptions. This breakthrough significantly enhances the intelligence and autonomy of UAVs in aerial scene understanding. This paper presents a comprehensive survey of OVOD in the context of UAV aerial scenes. We begin by aligning the core principles of OVOD with the unique characteristics of UAV vision, setting the stage for a specialized discussion. Building on this foundation, we construct a systematic taxonomy that categorizes existing OVOD methods for aerial imagery and provides a comprehensive overview of the relevant datasets. This structured review enables us to critically dissect the key challenges and open problems at the intersection of these fields. Finally, based on this analysis, we outline promising future research directions and application prospects. This survey aims to provide a clear road map and a valuable reference for both newcomers and seasoned researchers, fostering innovation in this rapidly evolving domain. We keep tracing related works at https://github.com/zhouyang2002/OVOD-in-UVA-imagery

无人机视觉开放词汇目标检测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。