arXiv:2501.07396cs.CV2025-01被引 3

用大模型零样本识别未知场景中的军事目标

Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models

  • 结合开放世界检测与大视觉语言模型,实现零样本目标识别
  • 在未知环境下对军事车辆识别准确率达78.3%
  • 适合军事、安防等需应对未知目标的高安全场景

自动目标识别(ATR)在导航和监视等任务中至关重要,安全性与准确性尤为关键。在军事等极端场景中,因未知地形、环境条件及新物体类别存在,现有检测器难以可靠工作。尽管大型视觉语言模型(LVLMs)具备零样本识别能力,但其定位性能不足。为此,本文提出一种新流程,融合开放世界检测器的定位能力与LVLMs的识别置信度,构建鲁棒的零样本ATR系统。研究对比了多种LVLM在军事车辆识别上的表现,发现模型在远距离(>500米)下仍可保持78.3%的识别准确率;并分析了距离范围、模态类型和提示方法对性能的影响,为未来在新类别与未知领域下的高可靠性ATR系统设计提供依据。

原文摘要 · Abstract (English)

Automatic target recognition (ATR) plays a critical role in tasks such as navigation and surveillance, where safety and accuracy are paramount. In extreme use cases, such as military applications, these factors are often challenged due to the presence of unknown terrains, environmental conditions, and novel object categories. Current object detectors, including open-world detectors, lack the ability to confidently recognize novel objects or operate in unknown environments, as they have not been exposed to these new conditions. However, Large Vision-Language Models (LVLMs) exhibit emergent properties that enable them to recognize objects in varying conditions in a zero-shot manner. Despite this, LVLMs struggle to localize objects effectively within a scene. To address these limitations, we propose a novel pipeline that combines the detection capabilities of open-world detectors with the recognition confidence of LVLMs, creating a robust system for zero-shot ATR of novel classes and unknown domains. In this study, we compare the performance of various LVLMs for recognizing military vehicles, which are often underrepresented in training datasets. Additionally, we examine the impact of factors such as distance range, modality, and prompting methods on the recognition performance, providing insights into the development of more reliable ATR systems for novel conditions and classes.

目标识别大模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。