用民用模型和数据集测试军用目标检测效果,发现仍需领域内训练。
Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition
- 直接使用现成模型与微调对比,评估其军用适用性。
- 大模型性能优于小模型,DETR表现接近甚至超越YOLO系列。
- 在空对地场景下,小目标检测仍是主要难点,适合军事视觉研究者。
自动目标检测与识别(ATD/R)对军事决策支持和(半)自主作战至关重要。尽管目标检测与人工智能的进步显著提升了ATD/R潜力,但公开军事数据集稀缺限制了系统应用。本文探索利用公开模型与民用数据集在军事场景中实现合理性能。我们在新获取的军事相关数据集上基准测试了六种YOLO迭代版本及两种DETR变体。该数据集包含军事车辆与复杂情形,如不同程度遮挡和小目标。评估采用原始模型与在VisDrone数据集上微调后的版本,该数据集具备小目标、空对地(A2G)视角和相关类别,可能泛化至军事任务。通过[email protected]与[email protected]:0.95指标,在A2G与地面对地面(G2G)视角、目标大小与模型规模下比较性能,揭示模型实时能力。主要发现:(1) 大模型优于小模型;(2) DETR类模型表现优于或媲美YOLO系列;(3) 在域外A2G数据集上微调可提升模型在A2G场景下的表现,并轻微改善小目标检测;(4) 所有模型在空对地场景中小目标检测仍表现不佳。结论:尽管目标检测技术进步,但领域内训练对构建有效ATD/R系统仍至关重要。
原文摘要 · Abstract (English)
Automatic Target Detection and Recognition (ATD/R) is critical for military decision support and (semi-)autonomous operations. Recent advances in object detection and artificial intelligence (AI) significantly boosted the potential performance of ATD/R. However, the scarcity of publicly available military datasets limits the application of these systems. As a solution, this paper explores the use of publicly available models and civilian datasets to achieve reasonable performance in military contexts. We benchmark several state-of-the-art models, including six iterations of the YOLO series and two variations on the DETR framework, on a newly acquired military relevant dataset. This dataset features military vehicles and challenging circumstances, including various degrees of occlusions and small targets. The out-of-the-box version of each model is validated alongside a version finetuned on the VisDrone dataset. This dataset features small objects, an Air-to-Ground (A2G) perspective and relevant classes, potentially generalizing to our military ATD/R task. We compare the performance of the models using [email protected] and [email protected]:0.95, across A2G and Ground-to-Ground (G2G) perspective, target size and model size, giving insight into the real-time capabilities of models. Our main findings are: (1) bigger models outperform smaller models, (2) DETR-based models show promising results compared to the YOLO series,(3) fine-tuning models on an out-of-domain A2G dataset, improves their A2G performance and slightly improves their performance on small objects, but (4) all models still struggle with detecting small objects in an A2G scenario. We conclude that, despite recent advances in object detection, in-domain training is still crucial for creating capable ATD/R systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。