arXiv:2411.03491cs.CV2024-11AAAI被引 1

用自然语言定义目标,实现无需训练的自动识别系统

An Application-Agnostic Automatic Target Recognition System Using Vision Language Models

  • 通过文本或图像示例动态定义目标类别,支持非技术人员使用
  • 结合帧序列信息实现目标轨迹匹配与框体重评分,提升检测精度
  • 生成热力图与拼接影像,直观呈现空中侦察区域的目标分布

我们提出一种新型自动目标识别(ATR)系统,基于开放词汇的物体检测与分类模型。该方法的核心优势在于:目标类别可在运行时由非技术用户通过少量自然语言描述、图像样例或两者结合的方式灵活定义,特别适用于缺乏训练数据的独特目标。系统引入多项创新技术,包括利用连续重叠帧中的额外信息进行管段识别(即序列边界框匹配)、边界框重评分以及管段连接。此外,我们开发了一种可视化方法,将多帧重叠输出合成一幅扫描区域的拼接图,并生成目标检测的核密度估计(热力图)。初始应用聚焦于机场跑道未爆弹药的探测与清除,目前正拓展至其他实际场景。

原文摘要 · Abstract (English)

We present a novel Automatic Target Recognition (ATR) system using open-vocabulary object detection and classification models. A primary advantage of this approach is that target classes can be defined just before runtime by a non-technical end user, using either a few natural language text descriptions of the target, or a few image exemplars, or both. Nuances in the desired targets can be expressed in natural language, which is useful for unique targets with little or no training data. We also implemented a novel combination of several techniques to improve performance, such as leveraging the additional information in the sequence of overlapping frames to perform tubelet identification (i.e., sequential bounding box matching), bounding box re-scoring, and tubelet linking. Additionally, we developed a technique to visualize the aggregate output of many overlapping frames as a mosaic of the area scanned during the aerial surveillance or reconnaissance, and a kernel density estimate (or heatmap) of the detected targets. We initially applied this ATR system to the use case of detecting and clearing unexploded ordinance on airfield runways and we are currently extending our research to other real-world applications.

目标识别视觉语言模型开放词汇无人机侦察

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。