arXiv:2503.00057cs.CVcs.CL2025-03被引 29

用大模型生成数据训练的YOLOv12在苹果检测中表现最佳。

Improved YOLOv12 with LLM-Generated Synthetic Data for Enhanced Apple Detection and Benchmarking Against YOLOv11 and YOLOv10

  • 用大语言模型生成合成图像训练YOLOv12,无需实地采集数据。
  • YOLOv12n在精度、召回率和mAP@50上均领先前代模型。
  • 适合农业自动化领域需要高精度目标检测的研究者使用。

本研究评估了YOLOv12在苹果检测中的性能,并与YOLOv11和YOLOv10进行对比,所有模型均在由大语言模型(LLMs)生成的合成图像上完成训练。YOLOv12n配置达到最高精度0.916,最高召回率0.969,以及最高mAP@50值0.978。相比之下,YOLOv11系列中表现最佳的是YOLOv11x,其精度为0.857,召回率为0.85,mAP@50为0.91;而YOLOv10系列中,YOLOv10b和YOLOv10l精度均为0.85,YOLOv10n召回率最高达0.8,mAP@50为0.89。结果表明,基于真实感合成数据训练的YOLOv12在关键指标上超越了前代模型。该方法显著降低农业场景下人工数据采集成本。此外,计算效率对比显示,YOLOv11n推理时间最短,为4.7毫秒,优于YOLOv12n的5.6毫秒和YOLOv10n的5.9毫秒。尽管YOLOv12在精度上领先于YOLOv11和YOLOv10,但当前仍由YOLOv11n保持最快推理速度。

原文摘要 · Abstract (English)

This study evaluated the performance of the YOLOv12 object detection model, and compared against the performances YOLOv11 and YOLOv10 for apple detection in commercial orchards based on the model training completed entirely on synthetic images generated by Large Language Models (LLMs). The YOLOv12n configuration achieved the highest precision at 0.916, the highest recall at 0.969, and the highest mean Average Precision (mAP@50) at 0.978. In comparison, the YOLOv11 series was led by YOLO11x, which achieved the highest precision at 0.857, recall at 0.85, and mAP@50 at 0.91. For the YOLOv10 series, YOLOv10b and YOLOv10l both achieved the highest precision at 0.85, with YOLOv10n achieving the highest recall at 0.8 and mAP@50 at 0.89. These findings demonstrated that YOLOv12, when trained on realistic LLM-generated datasets surpassed its predecessors in key performance metrics. The technique also offered a cost-effective solution by reducing the need for extensive manual data collection in the agricultural field. In addition, this study compared the computational efficiency of all versions of YOLOv12, v11 and v10, where YOLOv11n reported the lowest inference time at 4.7 ms, compared to YOLOv12n's 5.6 ms and YOLOv10n's 5.9 ms. Although YOLOv12 is new and more accurate than YOLOv11, and YOLOv10, YOLO11n still stays the fastest YOLO model among YOLOv10, YOLOv11 and YOLOv12 series of models. (Index: YOLOv12, YOLOv11, YOLOv10, YOLOv13, YOLOv14, YOLOv15, YOLOE, YOLO Object detection)

目标检测农业视觉合成数据YOLO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。