用大模型提升雷达图像军用车辆识别,准确率达98%。
Towards a Large Language-Vision Question Answering Model for MSTAR Automatic Target Recognition
- 用参数高效微调训练大模型处理合成孔径雷达图像
- 在MSTAR数据集上实现98%的细粒度目标识别准确率
- 为军事情报领域的遥感自动目标识别提供新思路
大型语言视觉模型(LLVM)如OpenAI的ChatGPT和GPT-4,在文本与图像分析中表现出强大能力。将文本与图像数据融合代表了自动目标识别(ATR)领域的重要范式转变。基于Transformer的最新研究显著提升了地理空间感知任务的表现。本研究聚焦于合成孔径雷达(SAR)图像的图像描述生成与视觉问答(VQA),考察了CLIP与LLaVA等前沿模型架构。我们基于MSTAR公开数据集构建了一个正在开发的SAR训练与评估基准,新增了描述性文本标注及问答对。该挑战数据集旨在推动大模型在复杂环境下的细微特征识别能力。通过参数高效微调,我们实现了98%的细粒度目标识别准确率。本文详细阐述了数据构建过程与实验设计,指出了可能导致误判的潜在陷阱。在复杂环境下准确识别并区分军用车辆类型是关键挑战,人类分析师需数月培训与多年经验。本研究首次系统性探索将大模型应用于SAR场景,推进了军事与情报背景下的机器辅助遥感自动目标识别。
原文摘要 · Abstract (English)
Large language-vision models (LLVM), such as OpenAI's ChatGPT and GPT-4, have gained prominence as powerful tools for analyzing text and imagery. The merging of these data domains represents a significant paradigm shift with far-reaching implications for automatic target recognition (ATR). Recent transformer-based LLVM research has shown substantial improvements for geospatial perception tasks. Our study examines the application of LLVM to remote sensing image captioning and visual question-answering (VQA), with a specific focus on synthetic aperture radar (SAR) imagery. We examine newly published LLVM methods, including CLIP and LLaVA neural network transformer architectures. We have developed a work-in-progress SAR training and evaluation benchmark derived from the MSTAR Public Dataset. This has been extended to include descriptive text captions and question-answer pairs for VQA tasks. This challenge dataset is designed to push the boundaries of an LLVM in identifying nuanced ATR details in SAR imagery. Utilizing parameter-efficient fine-tuning, we train an LLVM method to identify fine-grained target qualities at 98% accuracy. We detail our data setup and experiments, addressing potential pitfalls that could lead to misleading conclusions. Accurately identifying and differentiating military vehicle types in SAR data poses a critical challenge, especially under complex environmental conditions. Mastering this target recognition skill may require a human analyst months of training and years of practice. This research represents a unique effort to apply LLVM to SAR applications, advancing machine-assisted remote sensing ATR for military and intelligence contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。