arXiv:2502.13936cs.CVcs.LG2025-02中稿 · VISAPP 2025被引 5

图像合成比传统增强更有效,提升目标检测精度。

Image compositing is all you need for data augmentation

  • 用图像合成替代传统增强方法,提升模型泛化能力。
  • 在自定义飞机数据集上,[email protected] 提升显著,优于扩散模型。
  • 适合小样本场景下的目标检测模型优化,尤其军事/商用飞机识别。

本文研究了多种数据增强技术对目标检测模型性能的影响。以YOLOv8为基准,在包含商业与军用飞机的自定义数据集上进行微调,对比经典增强、图像合成及Stable Diffusion XL与ControlNet等生成模型的效果。实验表明,图像合成在精度、召回率和[email protected]指标上表现最佳,显著优于其他方法;生成模型也带来明显提升,证明先进增强技术对目标检测的潜力。结果强调数据集多样性与增强策略对真实场景泛化的重要性。未来工作将探索半监督学习融合与进一步优化,以应对更大更复杂数据集的挑战。

原文摘要 · Abstract (English)

This paper investigates the impact of various data augmentation techniques on the performance of object detection models. Specifically, we explore classical augmentation methods, image compositing, and advanced generative models such as Stable Diffusion XL and ControlNet. The objective of this work is to enhance model robustness and improve detection accuracy, particularly when working with limited annotated data. Using YOLOv8, we fine-tune the model on a custom dataset consisting of commercial and military aircraft, applying different augmentation strategies. Our experiments show that image compositing offers the highest improvement in detection performance, as measured by precision, recall, and mean Average Precision ([email protected]). Other methods, including Stable Diffusion XL and ControlNet, also demonstrate significant gains, highlighting the potential of advanced data augmentation techniques for object detection tasks. The results underline the importance of dataset diversity and augmentation in achieving better generalization and performance in real-world applications. Future work will explore the integration of semi-supervised learning methods and further optimizations to enhance model performance across larger and more complex datasets.

目标检测数据增强图像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。