用合成图像提升面包回收检测模型性能
Training a Computer Vision Model for Commercial Bakeries with Primarily Synthetic Images

- 用pix2pix和CycleGAN生成合成图像扩充数据集
- 最佳模型在测试集上达到90.3%的平均精度
- 适合工业质检与小样本视觉任务研究者
在食品行业,重新处理退货产品是提高资源效率的关键步骤。[SBB23]提出了一项AI应用,用于自动化追踪退回的面包卷。本文在此基础上构建了一个包含2432张图像、涵盖更广泛烘焙产品的扩展数据集。为增强模型鲁棒性,采用生成模型pix2pix和CycleGAN生成合成图像。我们在检测任务上训练了最先进的目标检测模型YOLOv8和YOLOv9。最终最佳模型在测试集上的平均精度([email protected])达到90.3%。
原文摘要 · Abstract (English)
In the food industry, reprocessing returned product is a vital step to increase resource efficiency. [SBB23] presented an AI application that automates the tracking of returned bread buns. We extend their work by creating an expanded dataset comprising 2432 images and a wider range of baked goods. To increase model robustness, we use generative models pix2pix and CycleGAN to create synthetic images. We train state-of-the-art object detection model YOLOv9 and YOLOv8 on our detection task. Our overall best-performing model achieved an average precision [email protected] of 90.3% on our test set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。