用可解释AI优化生成数据,提升红外图像车辆检测效果。
Improving Object Detection by Modifying Synthetic Data with Explainable AI
- 用XAI热力图指导3D模型修改,动态调节生成图像真实度。
- 在红外车辆检测中,检测性能提升至mAP50=96.1%。
- 无需大量人工干预,适合缺乏标注数据的场景。
真实世界数据有限严重制约计算机视觉模型性能,尤其对训练中罕见样本。合成图像虽有潜力,但如何设计最优合成数据仍不明确,且人工设计耗时费力。本文提出一种新方法:利用稳健的可解释AI(XAI)技术,指导人机协作修改用于生成图像的3D网格模型。该框架支持同时增加和减少图像真实度,均可提升模型表现。以红外图像中车辆检测为例,在ATR DSIAC数据集上,初始YOLOv8模型经合成数据微调后,对未见朝向车辆的检测准确率提升4.6%(mAP50达94.6%)。进一步通过XAI引导优化,性能再提升1.5%(达到96.1%),通过调整不同部分的真实度减少误分类。结果证明,该方法可实现精细化、可解释的合成数据定制,显著降低人工设计负担。
原文摘要 · Abstract (English)
Limited real-world data severely impacts model performance in many computer vision domains, particularly for samples that are underrepresented in training. Synthetically generated images are a promising solution, but 1) it remains unclear how to design synthetic training data to optimally improve model performance (e.g, whether and where to introduce more realism or more abstraction) and 2) the domain expertise, time and effort required from human operators for this design and optimisation process represents a major practical challenge. Here we propose a novel conceptual approach to improve the efficiency of designing synthetic images, by using robust Explainable AI (XAI) techniques to guide a human-in-the-loop process of modifying 3D mesh models used to generate these images. Importantly, this framework allows both modifications that increase and decrease realism in synthetic data, which can both improve model performance. We illustrate this concept using a real-world example where data are sparse; detection of vehicles in infrared imagery. We fine-tune an initial YOLOv8 model on the ATR DSIAC infrared dataset and synthetic images generated from 3D mesh models in the Unity gaming engine, and then use XAI saliency maps to guide modification of our Unity models. We show that synthetic data can improve detection of vehicles in orientations unseen in training by 4.6% (to mAP50 = 94.6%). We further improve performance by an additional 1.5% (to 96.1%) through our new XAI-guided approach, which reduces misclassifications through both increasing and decreasing the realism of different parts of the synthetic data. Our proof-of-concept results pave the way for fine, XAI-controlled curation of synthetic datasets tailored to improve object detection performance, whilst simultaneously reducing the burden on human operators in designing and optimising these datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。